Papers with natural language processing

300 papers
CogKTR: A Knowledge-Enhanced Text Representation Toolkit for Natural Language Understanding (2022.emnlp-demos)

Copied to clipboard

Challenge: Existing knowledge-enhanced methods are limited to knowledge-intensive tasks.
Approach: They propose a knowledge-enhanced text representation toolkit for natural language understanding . it combines knowledge acquisition, knowledge representation, knowledge injection and knowledge application .
Outcome: The proposed toolkit supports knowledge acquisition, knowledge representation, knowledge injection, and knowledge application.
AMesure: A Web Platform to Assist the Clear Writing of Administrative Texts (2020.aacl-demo)

Copied to clipboard

Challenge: OECD, 2016) report that a significant proportion of citizens still have general reading difficulties.
Approach: They propose to use a readability formula and natural language processing tools to analyze texts and highlight linguistic phenomena considered difficult to read.
Outcome: The AMesure platform analyzes administrative texts and offers advice from plain language guides.
Cross-lingual Semantic Representation for NLP with UCCA (2020.coling-tutorials)

Copied to clipboard

Challenge: introductory tutorial to UCCA, a symbolic meaning representation for semantic representations.
Approach: This tutorial introduces UCCA, a cross-linguistically applicable framework for semantic representation . it will provide a detailed introduction to the UCca annotation guidelines, design philosophy and available resources .
Outcome: The tutorial will provide a detailed introduction to the UCCA framework and compare it to other meaning representations.
Arabic Natural Language Processing (2022.emnlp-tutorials)

Copied to clipboard

Challenge: This tutorial provides background information for system developers and researchers working with Arabic in its various forms.
Approach: This tutorial provides the necessary background information for working with Arabic in its various forms.
Outcome: This tutorial will explain various Arabic linguistic phenomena and review the state-of-the-art in Arabic processing.
High Performance Natural Language Processing (2020.emnlp-tutorials)

Copied to clipboard

Challenge: a tutorial on scaling natural language processing will recapitulate the state-of-the-art in the field .
Approach: This cutting-edge tutorial recapitulates the state-of-the-art in natural language processing with scale in perspective.
Outcome: This cutting-edge tutorial recapitulates the state-of-the-art in natural language processing with scale in perspective.
TextPruner: A Model Pruning Toolkit for Pre-Trained Language Models (2022.acl-demo)

Copied to clipboard

Challenge: Large pre-trained language models have been used for many NLP tasks but computational resources are limited.
Approach: They propose an open-source model pruning toolkit for pre-trained language models . they propose a self-supervised pruning method that can be applied without labeled data.
Outcome: The proposed pruning method reduces model size without retraining the model and speeds up inference speed on the common CPU and GPU devices.
Beyond Multiword Expressions: Processing Idioms and Metaphors (P18-5)

Copied to clipboard

Challenge: idioms and metaphors processing is a rapidly growing area in NLP, says dr. s. robertson . idiomatic idiomas are characteristic to all areas of human activity and to all types of discourse.
Approach: This tutorial will provide attendees with a clear notion of idioms and metaphors . it will provide them with computational models of linguistic characteristics and methods .
Outcome: This tutorial aims to provide attendees with a clear notion of the linguistic characteristics of idioms and metaphors . it outlines how to model idiomatic idiomes and their processing and what resources are available to support their use .
A Benchmark Suite of Japanese Natural Questions (2024.starsem-1)

Copied to clipboard

Challenge: Existing studies to solve QA tasks in an integrated manner are not available in other languages because of the lack of QA datasets.
Approach: They build a Japanese version of Natural Questions using natural questions from query logs of a search engine and crowdsource it using crowdsourcing.
Outcome: The proposed datasets are based on natural questions from Japanese search engines and crowdsourced.
Big AI is Accelerating the Metacrisis: What Can We Do? (2026.acl-short)

Copied to clipboard

Challenge: LLM engineering is at the core of the problem of ecological, meaning, and language crises . big AI is fueling global crises and creating wealth and power for a handful of individuals and corporations while causing existential harm to life on earth.
Approach: et al., 2025, p162ff) argue that big AI is escalating global crises and creating a metacrisis.
Outcome: the field of natural language processing is at the core of the problem . it is being leveraged to create unprecedented wealth and power for a handful of individuals and corporations while causing existential harm to life on earth.
Network Features Based Co-hyponymy Detection (L18-1)

Copied to clipboard

Challenge: Existing methods to detect lexical relations have been used to identify them in both supervised and unsupervised ways.
Approach: They propose to use distributional semantic models to detect co-hyponymy relation with high accuracy and various network measures to perform better or at par with the state-of-the-art models.
Outcome: The proposed model performs better or at par with the state-of-the-art models.
pNLP-Mixer: an Efficient all-MLP Architecture for Language (2023.acl-industry)

Copied to clipboard

Challenge: large pre-trained language models are impractical for on-device applications due to their size and inference cost.
Approach: They propose an embedding-free MLP-Mixer model for on-device NLP that achieves high weight-efficiency thanks to a novel projection layer.
Outcome: The proposed model beats state-of-the-art of tiny models by 97.8% on two datasets . it beats mBERT on MTOP and multiATIS, while using 170x less parameters .
Synthetic Data in the Era of Large Language Models (2025.acl-tutorials)

Copied to clipboard

Challenge: 'synthetic data' is a data generated with the assistance of large language models to make dataset construction faster and cheaper.
Approach: This tutorial seeks to build a shared understanding of recent progress in synthetic data generation from NLP and related fields by grouping and describing major methods, applications, and open problems.
Outcome: This tutorial will describe methods, applications, and open problems that have been developed and are being used to improve the quality and efficiency of synthetic data generation.
Enhancing Self-Attention with Knowledge-Assisted Attention Maps (2022.naacl-main)

Copied to clipboard

Challenge: Existing works of knowledge infusion depend on multi-task learning frameworks, which are inefficient and require large-scale retraining when new knowledge is considered.
Approach: They propose a method which integrates knowledge-generated attention maps into the self-attention mechanism and integrates it into the model.
Outcome: The proposed model outperforms existing methods on academic datasets and industry-scale ad relevance applications.
OpenPrompt: An Open-source Framework for Prompt-learning (2022.acl-demo)

Copied to clipboard

Challenge: Prompt-learning is a new paradigm in natural language processing, adapting pre-trained language models to cloze-style prediction, autoregressive modeling, or sequence to sequence generation.
Approach: They propose a framework for prompt-learning that integrates pre-trained language models with a unified framework.
Outcome: The proposed framework is easy to use and flexible enough to integrate with other frameworks.
CULG: Commercial Universal Language Generation (2022.naacl-industry)

Copied to clipboard

Challenge: Pre-trained language models have improved performance for many NLP tasks in finance and healthcare.
Approach: They propose a large-scale commercial universal language generation model which is pre-trained on a corpus drawn from 10 markets across 7 languages.
Outcome: The proposed model outperforms other models on commercial generation tasks and on other markets, languages, and tasks.
LightSeq: A High Performance Inference Library for Transformers (2021.naacl-industry)

Copied to clipboard

Challenge: Existing inference frameworks for natural language processing are not the best choice for online service of sequence processing problems.
Approach: They propose a highly efficient inference library for Transformer models that includes GPU optimization techniques to streamline computation and reduce memory footprint.
Outcome: The proposed library achieves 14x speedup compared with TensorFlow and 1.4x speed up compared to a concurrent CUDA implementation.
Decoding Brain Activity Associated with Literal and Metaphoric Sentence Comprehension Using Distributional Semantic Models (2020.tacl-1)

Copied to clipboard

Challenge: Existing research has focused on applying semantic models to decode brain activity associated with the meaning of individual words.
Approach: They evaluate a range of semantic models to capture metaphor processing in the brain . they found that compositional models and word embeddings capture differences in the processing of literal and metaphoric sentences .
Outcome: The proposed models capture differences in the processing of literal and metaphoric sentences, providing support for the idea that the literal meaning is not fully accessible during familiar metaphor comprehension.
Mining Tweets that refer to TV programs with Deep Neural Networks (D19-55)

Copied to clipboard

Challenge: opinion mining is a popular natural language processing technique, but a problem is robustness for user-generated texts . a recent study shows that a model that handles context can extract the opinion target with 90% accuracy .
Approach: They propose a model that handles context in many natural language processing areas to solve a problem of extracting opinion references from text.
Outcome: Experiments on tweets that refer to television programs show the proposed model can extract opinion references with more than 90% accuracy.
Development of Conversational AI for Sleep Coaching Programme (2021.eacl-srw)

Copied to clipboard

Challenge: Existing methods to treat insomnia neglect conversational aspects, which plays a critical role in sleep therapy.
Approach: They propose to develop conversational AI for a sleep coaching programme which is motivated by CBT-I treatment and provide an automated analytic system to support human experts.
Outcome: The proposed system could interact naturally with a user and provide an automated analytic system to support human experts.
Newspaper Signaling for Crisis Prediction (2024.naacl-demo)

Copied to clipboard

Challenge: Existing systems for detecting crisis-related signals are limited due to unstructured data, media, and cultural bias, and multiple languages.
Approach: They propose a model for multi-lingual and open-domain newspaper signaling for detecting crisis-related indicators in newspaper articles.
Outcome: The proposed model can detect crisis-related indicators in multiple languages and can be used in open crisis domains in real-time.
Friend-training: Learning from Models of Different but Related Tasks (2023.eacl-main)

Copied to clipboard

Challenge: Current self-training methods focus on improving model performance on a single task.
Approach: They propose a cross-task self-training framework where models trained to do different tasks are used in iterative training, pseudo-labeling, and retraining processes to help each other for better selection of pseudo-labeled labels.
Outcome: The proposed framework achieves the best performance compared to baselines on two dialogue understanding tasks.
Code Representation Pre-training with Complements from Program Executions (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing languages have syntactic representations of code to improve code intelligence, but they are difficult to learn from code.
Approach: They propose to embed dynamic information of programs revealed by their test cases into feature representations of code as complements.
Outcome: The proposed method yields 6%/19% mAP improvements over its masked language modeling counterparts.
Regularized Graph Convolutional Networks for Short Text Classification (2020.coling-industry)

Copied to clipboard

Challenge: Short text classification is a problem in natural language processing, social network analysis, and e-commerce.
Approach: They propose a short text classification technique that incorporates label dependencies into the output space to overcome the limitations of short text.
Outcome: The proposed model outperforms baseline methods on proprietary and external datasets and is more robust to noise in textual features.
Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for hyperbole and metaphor detection focus on superficial text features, ignoring the associations of hyperbola and metaphor . Existing frameworks focus on identifying superficial text, focusing on superficial features .
Approach: They propose an emotion-guided hyperbole and metaphor detection framework based on bidirectional dynamic interaction.
Outcome: The proposed framework outperforms baseline methods on four datasets.
Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages (2025.acl-srw)

Copied to clipboard

Challenge: Low-resource languages (LRLs) face significant challenges in natural language processing due to limited data.
Approach: They evaluate adapter-based methods for adapting mLMs to low-resource languages . they use unstructured text and structured knowledge from ConceptNet to evaluate adapters .
Outcome: The proposed methods outperform large language models and LLaMA-3 and deepSeek-R1 models on low training data.
Design of BCCWJ-EEG: Balanced Corpus with Human Electroencephalography (2020.lrec-1)

Copied to clipboard

Challenge: Recent research has focused on the fusion of NLP and neuroscience of language.
Approach: They propose to use a balanced corpus of written Japanese (BCCWJ) annotated with human electroencephalography to improve annotations and annotations.
Outcome: The proposed language resource is annotated with human electroencephalography (EEG) and can improve on annotations, genres, languages, etc.
LOT: A Story-Centric Benchmark for Evaluating Chinese Long Text Understanding and Generation (2022.tacl-1)

Copied to clipboard

Challenge: Existing benchmarks for natural language processing focus on understanding or generating short texts . lack of standardized benchmarks makes it difficult to assess and compare models .
Approach: They propose a story-centric benchmark for Chinese long text modeling that aggregates two understanding tasks and two generation tasks.
Outcome: The proposed model outperforms similar-sized models on understanding and generation tasks.
Walia-LLM: Enhancing Amharic-LLaMA by Integrating Task-Specific and Generative Datasets (2024.findings-emnlp)

Copied to clipboard

Challenge: Low-resource languages are left behind due to the unavailability of resources.
Approach: They propose to integrate task-specific and generative datasets to improve language model performance for Amharic by fine-tuning an Amharican instruction fine-to-tuned model.
Outcome: The proposed model shows promising results in different NLP tasks and compares translated instruction datasets with the original model.
Fréchet Distance for Offline Evaluation of Information Retrieval Systems with Sparse Labels (2024.eacl-long)

Copied to clipboard

Challenge: Obtaining high-quality labeled data that accurately represents complexity of real-world scenarios can be expensive, time-consuming, or even impractical.
Approach: They propose to use Fréchet Inception Distance to measure distance between judged items and retrieved results.
Outcome: The proposed method improves on a MS MARCO dataset and TREC Deep Learning Tracks query sets.
Neural Networks in a Product of Hyperbolic Spaces (2022.naacl-srw)

Copied to clipboard

Challenge: Recent advances in the use of hyperbolic spaces have been reported in natural language processing and graph embedding.
Approach: They propose to extend hyperbolic neural networks to a product of hyperbolical spaces by using a single hyperbolically spaced hyperbole.
Outcome: The proposed method improves graph node classification accuracy on tree-like datasets.
Langsmith: An Interactive Academic Text Revision System (2020.emnlp-demos)

Copied to clipboard

Challenge: Currently, diversity and inclusion initiatives in the academic community are encouraged . however, writing papers in English can be a daunting task .
Approach: They propose a system that helps non-native English speakers to write papers in English . the system can suggest fluent, academic-style sentences based on their rough, incomplete phrases or sentences .
Outcome: The proposed system can help non-native English speakers write papers in English . the system can suggest fluent, academic-style sentences based on their rough sentences .
HERB: Measuring Hierarchical Regional Bias in Pre-trained Language Models (2022.findings-aacl)

Copied to clipboard

Challenge: Existing methods do not examine social groups categorised by geographical information, leaving the region-related biases in pre-trained LMs unexplored.
Approach: They propose a hierarchical regional bias evaluation method to quantify regional bias in pre-trained language models.
Outcome: The proposed method evaluates regional bias with regard to comprehensive topics and measures potential regional bias that can be propagated to downstream tasks.
Parameter-free and Accessible Prompt Learning to Enhance Adversarial Robustness for Pre-trained Vision-Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large pre-trained Vision-Language Models (VLMs) have revolutionized downstream vision-language tasks including classification, object detection, and segmentation.
Approach: They propose to search for text prompts at the word level rather than optimizing continuous textual embeddings to boost adversarial robustness.
Outcome: Experiments show that the proposed method outperforms hand-engineered prompts with average gains of +4.9% and +5.8%.
A Boundary-aware Neural Model for Nested Named Entity Recognition (D19-1)

Copied to clipboard

Challenge: Existing methods for named entity recognition ignore nested entities . a boundary-aware neural model can locate entities precisely by detecting boundaries .
Approach: They propose a boundary-aware neural model for nested named entity recognition which leverages entity boundaries to predict entity categorical labels.
Outcome: The proposed model outperforms state-of-the-art methods on GENIA dataset . it captures dependencies of entity boundaries and categorical labels, which helps to improve identifying entities.
Harmless Transfer Learning for Item Embeddings (2022.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to learn item embeddings for categorical features are limited by the frequency of items in real-world.
Approach: They propose a method that transfers knowledge from frequent items to rare items by introducing an auxiliary transfer loss.
Outcome: The proposed framework significantly boosts the performance on a variety of NLP and recommendation system tasks.
Grokking of Hierarchical Structure in Vanilla Transformers (2023.acl-short)

Copied to clipboard

Challenge: a recent study has shown that neural sequence models like transformers can generalize hierarchically when training for extended periods.
Approach: They show that transformers can learn to generalize hierarchically after long training periods . they call this phenomenon structural grokking, which exhibits inverted U-shaped scaling in model depth .
Outcome: The proposed model generalizes better than both very deep and very shallow models on multiple datasets.
BeamR: Beam Reweighing with Attribute Discriminators for Controllable Text Generation (2022.findings-aacl)

Copied to clipboard

Challenge: Recent advances in natural language processing have led to the availability of large pre-trained language models with rich generative capabilities.
Approach: They propose a method to combine generative LMs with attribute discriminators to control different attributes of text generation.
Outcome: The proposed method performs better than existing state-of-the-art approaches in sentiment steering and machine translation formality tasks.
Can Post-Training Quantization Benefit from an Additional QLoRA Integration? (2025.naacl-industry)

Copied to clipboard

Challenge: Large language models require considerable computing resources, which can be costly and often unavailable.
Approach: They propose to integrate 4-bit Post-training Quantization with QLoRA to address these issues . they demonstrate that the integration outperforms standard quantization and fine-tuning .
Outcome: The proposed integration outperforms standard PTQ and 16-bit full-parameter fine-tuning on LLMs.
STAR: Spectral Truncation and Rescale for Model Merging (2025.naacl-short)

Copied to clipboard

Challenge: Model merging is an efficient way of obtaining a multi-task model from several pretrained models without further fine-tuning.
Approach: They propose a model merging technique that aims at mitigating "merging conflicts" by truncating small components in the respective spectral spaces and then an automatic parameter rescaling scheme to retain the nuclear norm of the original matrix.
Outcome: The proposed model outperforms baseline models on flan-T5 by 4.2% and is robust to hyperparamater choice.
Task-driven Layerwise Additive Activation Intervention (2025.naacl-short)

Copied to clipboard

Challenge: Existing approaches to task adaptation rely heavily on heuristic rules or prompt inputs.
Approach: They propose a layer-wise additive activation intervention framework that steers the LMs’ generation process by identifying and manipulating the activations.
Outcome: The proposed framework improves the accuracy of pretrained LMs and competing baselines on various datasets, demonstrating improvements in the accuracy and sample efficiency of the proposed framework.
PaperMage: A Unified Toolkit for Processing, Representing, and Manipulating Visually-Rich Scientific Documents (2023.emnlp-demo)

Copied to clipboard

Challenge: Existing tools for working with scientific documents are limited and documents are often in difficult-to-use PDF formats.
Approach: They propose an open-source Python toolkit for analyzing and processing visually-rich scientific documents.
Outcome: PaperMage provides turn-key recipes for common scientific document processing use-cases.
A Decade of Knowledge Graphs in Natural Language Processing: A Survey (2022.aacl-main)

Copied to clipboard

Challenge: Knowledge graphs (KGs) are a representation of semantic relations between entities . despite their popularity, there is still no general understanding of what exactly a KG is or for what tasks it is applicable.
Approach: They analyze 507 papers on knowledge graphs in natural language processing (NLP) they provide a taxonomy of tasks and review the maturity of individual research streams .
Outcome: The findings summarize the literature and highlight directions for future work.
A Computational Framework to Identify Self-Aspects in Text (2025.acl-srw)

Copied to clipboard

Challenge: a Ph.D. proposal aims to identify Self-aspects in text, which are underexplored in natural language processing . many aspects of the Self align with psychological and other well-researched phenomena .
Approach: They propose to develop a computational framework to identify Self-aspects in text . they will use an ontology of Self-facets and an annotated gold-standard dataset .
Outcome: The proposed framework will evaluate discriminative models, generative large language models, embedding-based retrieval approaches against four main criteria: interpretability, ground-truth adherence, accuracy, and computational efficiency.
Implicitly Abusive Language – What does it actually look like and why are we not getting there? (2021.naacl-main)

Copied to clipboard

Challenge: Existing datasets make learning implicit abuse difficult, argues a new position paper . a lack of work on implicit abuse has limited the effectiveness of automatic detection .
Approach: They argue that existing datasets make learning implicit abuse difficult . they propose a divide-and-conquer strategy to detect implicit abuse .
Outcome: The proposed model could be improved to detect implicit abuse in a dataset with a standardized model.
Swift Cross-Dataset Pruning: Enhancing Fine-Tuning Efficiency in Natural Language Understanding (2025.coling-main)

Copied to clipboard

Challenge: Current approaches for fine-tuning datasets rely on expensive sample ranking processes . data set pruning aims to select a subset of a dataset for efficient model training .
Approach: They propose a method that uses TF-IDF embeddings with geometric median to rapidly evaluate sample importance.
Outcome: The proposed method significantly reduces training and storage costs while maintaining model effectiveness.
Post-Abstention: Towards Reliably Re-Attempting the Abstained Instances in QA (2023.acl-long)

Copied to clipboard

Challenge: Despite remarkable progress made in natural language processing, even the state-of-the-art systems often make incorrect predictions.
Approach: They propose to use selective prediction to enable models to abstain from answering when their predictions are likely to be incorrect.
Outcome: The proposed method improves performance on 11 QA datasets and in- and out-of-domain settings.
Negation Detection in Dutch Spoken Human-Computer Conversations (2022.lrec-1)

Copied to clipboard

Challenge: Existing negation detection methods in English are not available.
Approach: They propose to annotate a Dutch dialogue corpus with negation cues and their scopes.
Outcome: The proposed method can detect negation cues and scope in Dutch dialogues with high precision and recall.
GTA: Supervised-Guided Reinforcement Learning for Text Classification with Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Reinforcement learning fine-tuning methods suffer from inefficient exploration and slow convergence . supervised fine- tuning methods have limited performance ceiling and less solid theoretical foundation .
Approach: They propose a Guess-Think-Answer framework that combines supervised and supervised learning in a unified training paradigm.
Outcome: The proposed framework outperforms both standalone SFT and RL training models on three text classification benchmarks.
Does GPT-3 Generate Empathetic Dialogues? A Novel In-Context Example Selection Method and Automatic Evaluation Metric for Empathetic Dialogue Generation (2022.coling-1)

Copied to clipboard

Challenge: Empathy is a multi-dimensional concept consisting of cognitive and affective aspects.
Approach: They propose two new in-context example selection methods that utilize emotion and situational information.
Outcome: The proposed method is effective in measuring the degree of human empathy.
Revisiting and Advancing Chinese Natural Language Understanding with Accelerated Heterogeneous Knowledge Pre-training (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing knowledge-enhanced pre-trained language models (KEPLMs) can capture internal knowledge, but can't understand external background knowledge.
Approach: They propose to use Chinese knowledge-enhanced pre-trained language models to improve context-aware representations via learning from structured relations in knowledge bases.
Outcome: Experiments show that Chinese knowledge-enhanced pre-trained language models outperform strong baselines over various benchmark NLP tasks and in different model sizes.
DropMix: A Textual Data Augmentation Combining Dropout with Mixup (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods to overcome overfitting in text learning do not consider dimensionality . dimensionalization is important for deep neural networks to overcome the problem .
Approach: They propose a saliency map-based approach to overcome overfitting in text learning . they propose augmentation regularization methods such as Dropout and Mixup to improve regularization .
Outcome: Empirical results show that the proposed approach overcomes overfitting in text learning . dropout and mixup methods are effective in enhancing regularization .
Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning (2024.acl-long)

Copied to clipboard

Challenge: Experimental results show that fine-tuning of large language models for specific tasks can be challenging . distribution shift during fine-timing can lead to performance degradation in general task capabilities .
Approach: They propose a new approach that bridges the distribution gap between task datasets and LLMs by guiding fine-tuning with a distilled dataset generated by the model itself.
Outcome: The proposed approach achieves comparable or superior performance on downstream tasks compared to the vanilla approach.
Guidance-Based Prompt Data Augmentation in Specialized Domains for Named Entity Recognition (2024.acl-short)

Copied to clipboard

Challenge: specialized fields such as science and biology face significant challenges due to the scarcity of quality data.
Approach: They propose a guidance data augmentation technique that abstracts context and sentence structure and maintains context-entity relationships for DA.
Outcome: The proposed method enhances the training performance of named entity recognition tasks while maintaining context-entity relationships.
A Survey in Automatic Irony Processing: Linguistic, Cognitive, and Multi-X Perspectives (2022.coling-1)

Copied to clipboard

Challenge: figurative language research has focused on sarcasm and irony, but there is still a gap in the field.
Approach: They propose to review computational irony, cognitive science, and neural models of irony processing . they aim to encourage a balanced and equal research environment in figurative languages .
Outcome: The proposed multi-X irony processing perspectives will provide an overview of computational irony, insights from linguisic theory and cognitive science, and interactions with downstream NLP tasks.
Neural Topic Modeling with Large Language Models in the Loop (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated promising capabilities in topic discovery, but their direct application to topic modeling suffers from issues such as incomplete topic coverage, misalignment of topics, and inefficiency.
Approach: They propose a novel LLM-in-the-loop framework that integrates Large Language Models with Neural Topic Models (NTMs) global topics and document representations are learned through the NTM, while an LLM refines these topics using an Optimal Transport (OT)-based alignment objective.
Outcome: The proposed framework improves topic interpretability while preserving the efficiency of existing NTMs.
Probing Across Time: What Does RoBERTa Know and When? (2021.findings-emnlp)

Copied to clipboard

Challenge: Current approaches to natural language processing rely on fixed artifacts such as language models . current studies have focused on how these models acquire and demonstrate knowledge .
Approach: They apply probing techniques to examine how language models acquire knowledge . they aim to inform future work on more efficient pretraining and understanding dependencies .
Outcome: The proposed model learns linguistic abstractions, factual and commonsense knowledge, and reasoning abilities fast, stably, and robustly across domains.
Order-Based Pre-training Strategies for Procedural Text Understanding (2024.naacl-short)

Copied to clipboard

Challenge: Procedural text is difficult to understand due to the changing attributes of entities in the context.
Approach: They propose sequence-based pre-training methods to enhance procedural understanding in natural language processing by using ordered instructions to guide individuals through a task.
Outcome: The proposed methods improve on two datasets in the datasets NPN-Cooking and ProPara domains respectively.
CoXQL: A Dataset for Parsing Explanation Requests in Conversational XAI Systems (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing systems based on large language models (LLMs) are more precise and reliable in identifying users’ intentions, but the recognition of intents still presents a challenge in the case of ConvXAI, since little training data exist and the domain is highly specific.
Approach: They propose to use a dataset in the NLP domain for user intent recognition in ConvXAI to improve parsing performance.
Outcome: The proposed system outperforms existing methods and improves on existing ones.
BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation (2025.coling-main)

Copied to clipboard

Challenge: Knowledge distillation (KD) is a method for reducing model size while preserving performance.
Approach: They propose a method to distill large language models at the logit level by transferring knowledge from a large teacher model to a smaller student model.
Outcome: The proposed method outperforms supervised fine-tuning, vanilla KL loss and five other distillation methods on 13 datasets.
Scientia Potentia Est—On the Role of Knowledge in Computational Argumentation (2022.tacl-1)

Copied to clipboard

Challenge: Existing research on argumentation models does not provide a systematic overview of the types of knowledge required in CA tasks.
Approach: They propose a taxonomy of the types of knowledge required in CA tasks . authors propose exploitation of these knowledge types for four main research areas .
Outcome: The proposed taxonomy proposes a systematic overview of the types of knowledge required in CA tasks.
Answerable or Not: Devising a Dataset for Extending Machine Reading Comprehension (C18-1)

Copied to clipboard

Challenge: Existing MRC algorithms assume that each question is answerable by looking at text passages, but to realize human-like language comprehension ability, a machine should be able to distinguish not-answerable questions from answerable questions.
Approach: They propose a method for automatically assigning difficulty level labels to a dataset that alters an existing MRC dataset and describes the resulting dataset.
Outcome: The proposed method can detect NAQs in a dataset with difficulty level labels and is valid and potentially useful in the development of advanced MRC models.
The Art of Abstention: Selective Prediction and Error Regularization for Natural Language Processing (2021.acl-long)

Copied to clipboard

Challenge: Pre-trained language models have improved the state-of-the-art results on many NLP applications.
Approach: They propose a simple error regularization trick that improves confidence estimation without substantially increasing the computation budget.
Outcome: The proposed regularization improves confidence estimation without increasing computation budget.
Is ChatGPT a General-Purpose Natural Language Processing Task Solver? (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in scale have enabled large language models to perform NLP tasks zero-shot . however, it is not known whether ChatGPT can serve as a generalist model that can perform many NLP jobs zero- shot.
Approach: They empirically evaluate ChatGPT's zero-shot learning ability on 20 popular NLP datasets . they find it performs well on many tasks favoring reasoning abilities .
Outcome: The proposed model can perform many NLP tasks zero-shot without adaptation on downstream data.
Robust Multilingual Part-of-Speech Tagging via Adversarial Training (N18-1)

Copied to clipboard

Challenge: Adversarial training (AT) is a powerful regularization method for neural networks, aiming to achieve robustness to input perturbations.
Approach: They propose and analyze a neural POS tagging model that exploits adversarial training by training on unmodified and adversarials.
Outcome: The proposed model improves overall tagging accuracy and prevents over-fitting in low resource languages and boosts tabbing accuracy for rare / unseen words.
Dealing with Controversy: An Emotion and Coping Strategy Corpus Based on Role Playing (2024.findings-emnlp)

Copied to clipboard

Challenge: Psychological studies aim at explaining internal mechanisms of emotions, while computational studies simplify them into labels.
Approach: They propose to treat emotions as strategies to cope with salient situations . they introduce a task of coping identification and a corpus constructed via role-playing .
Outcome: The proposed method allows to investigate the link between emotions and behavior, which also emerges in language.
Exploring the Role of Prior Beliefs for Argument Persuasion (N18-1)

Copied to clipboard

Challenge: Recent studies in natural language processing (NLP) have shown that the language of opinion holders and their patterns of interaction play a key role in changing the mind of a reader.
Approach: They propose to use a dataset to study the effect of language use vs. prior beliefs on persuasion in a controlled setting that takes into account political and religious ideology.
Outcome: The proposed controlled setting takes into account political and religious ideology and shows that prior beliefs play a more important role than language use effects.
Composite Backdoor Attacks Against Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated superior performance on various tasks, but untrustworthy third-party LLMs may covertly introduce vulnerabilities for downstream tasks.
Approach: They propose a composite backdoor attack that scatters multiple trigger keys in different prompt components.
Outcome: The proposed attack achieves 100% Attack Success Rate (ASR) with a False Triggered Rate (FTR) below 2.06% and negligible model accuracy degradation.
Cross-domain NER with Generated Task-Oriented Knowledge: An Empirical Study from Information Density Perspective (2024.emnlp-main)

Copied to clipboard

Challenge: Cross-domain Named Entity Recognition (CDNER) is crucial for Knowledge Graph (KG) construction and natural language processing (NLP)
Approach: They propose to automatically generate task-oriented knowledge using large language models (LLMs) and then employ task-orientated pre-training (TOPT) to facilitate domain adaptation.
Outcome: The proposed model can learn to distinguish between different entities and improve its domain adaptation.
MenatQA: A New Dataset for Testing the Temporal Comprehension and Reasoning Abilities of Large Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have shown nearly saturated performance on many NLP tasks.
Approach: They construct multiple sensitive factors time QA which encompasses three temporal factors . they test current mainstream LLMs with different parameter sizes .
Outcome: The proposed model incorporates three temporal factors with 2,853 samples . the results show that LLMs fall behind smaller models on these factors .
ConTextING: Granting Document-Wise Contextual Embeddings to Graph Neural Networks for Inductive Text Classification (2022.coling-1)

Copied to clipboard

Challenge: Graph neural networks (GNNs) are used to learn document representation from graph structures.
Approach: They propose a unified model with a joint training mechanism to learn from document embeddings and contextual word interactions simultaneously.
Outcome: The proposed model outperforms pure inductive GNNs and BERT-style models . the proposed model also has a joint training mechanism to learn from document embeddings and contextual word interactions simultaneously.
AHP-Powered LLM Reasoning for Multi-Criteria Evaluation of Open-Ended Responses (2024.findings-emnlp)

Copied to clipboard

Challenge: Question answering (QA) tasks have been extensively studied in the field of natural language processing.
Approach: They propose a method that leverages large language models and the analytic hierarchy process to assess open-ended questions.
Outcome: The proposed method more closely aligns with human judgment compared to baselines on four datasets.
Filling Missing Paths: Modeling Co-occurrences of Word Pairs and Dependency Paths for Recognizing Lexical Semantic Relations (N18-1)

Copied to clipboard

Challenge: Existing approaches to recognize lexical semantic relations between word pairs require that word pairs co-occur in a sentence.
Approach: They propose to exploit lexico-syntactic paths between two target words to exploit the semantic relations between word pairs.
Outcome: The proposed model can generalize the co-occurrences of word pairs and dependency paths and extract features capturing relational information from word pairs.
Label Anchored Contrastive Learning for Language Understanding (2022.naacl-main)

Copied to clipboard

Challenge: a novel approach to contrastive learning for language understanding is not fully explored . contrastive training has been widely applied to self-supervised representation learning .
Approach: They propose a label anchored contrastive learning approach for language understanding using a class label.
Outcome: The proposed approach improves on GLUE and CLUE benchmarks by 4.1% compared to the state-of-the-art approaches . the proposed approach also improves under the few-shot and data imbalance settings .
DPTDR: Deep Prompt Tuning for Dense Passage Retrieval (2022.coling-1)

Copied to clipboard

Challenge: Recent studies show that prompt tuning is unfriendly for industrial deployment in dense retrieval tasks.
Approach: They propose to apply prompt tuning to dense retrieval tasks to reduce deployment cost . they propose to use retrieval-oriented intermediate pretraining and unified negative mining .
Outcome: The proposed method outperforms state-of-the-art models on MS-MARCO and Natural Questions.
A Comprehensive Survey of Sentence Representations: From the BERT Epoch to the CHATGPT Era and Beyond (2024.eacl-long)

Copied to clipboard

Challenge: Sentence representations are a critical component in NLP applications such as retrieval, question answering, and text classification.
Approach: They present a systematic review of the literature on sentence representations focusing mostly on deep learning models.
Outcome: The proposed methods highlight the key contributions and challenges in this area and suggest potential avenues for improving the quality and efficiency of sentence representations.
Prefix Lexicalization of Synchronous CFGs using Synchronous TAG (P18-1)

Copied to clipboard

Challenge: epsilon-free, chain-free synchronous context-free grammars can be converted into weakly equivalent synchronous tree-adjoining grammars (STAGs) this transformation doubles the grammar’s rank and cubes its size, but in practice the size increase is only quadratic.
Approach: They extend Greibach normal form from CFGs to SCFGs and prove new formal properties about SCFG, a formalism with many applications in natural language processing.
Outcome: The proposed grammars achieve asymptotic and empirical speed improvements on a machine translation task.
EDTC: A Corpus for Discourse-Level Topic Chain Parsing (2021.findings-emnlp)

Copied to clipboard

Challenge: Discourse analysis is a fundamental part of natural language processing.
Approach: They propose a discourse-level topic chain parsing system which can be automated . they propose lexical cohesion modeling instead of lexically measuring topic structure .
Outcome: The proposed system is robust and reliable, and can provide high reliability and low confidence scores.
Exploring Data Augmentation for Code Generation Tasks (2023.findings-eacl)

Copied to clipboard

Challenge: Recent advances in natural language processing have impacted how models are trained for programming language tasks.
Approach: They propose to use augmentation methods that yield consistent improvements in code translation and summarization by up to 6.9% and 7.5% respectively.
Outcome: The proposed methods improve translation and summarization by 6.9% and 7.5% respectively.
Adapter Pruning using Tropical Characterization (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on adapter pruning have not examined the optimal number of adapter parameters needed for downstream applications.
Approach: They propose an adapter pruning approach that prunes adapter parameters without changing the orientation of underlying tropical hypersurfaces.
Outcome: The proposed approach prunes adapter layers without changing the orientation of underlying tropical hypersurfaces.
Smaller Text Classifiers with Discriminative Cluster Embeddings (N18-2)

Copied to clipboard

Challenge: Word embeddings dominate overall model sizes in neural methods for natural language processing, especially when large vocabularies and high dimensions are used.
Approach: They propose a Gumbel-Softmax distribution to maximize over the latent clustering while minimizing the task loss.
Outcome: The proposed method minimizes the task loss while maximizing over the latent clustering while remaining parameter-efficient.
Implicit n-grams Induced by Recurrence (2022.naacl-main)

Copied to clipboard

Challenge: Recent studies show that self-attention based models have limitations on modeling sequential transformations.
Approach: They propose to extract some explainable features from trained RNNs that are reminiscent of classical n-grams features.
Outcome: The proposed models can model interesting linguistic phenomena such as negation and intensification.
DuReader_robust: A Chinese Dataset Towards Evaluating Robustness and Generalization of Machine Reading Comprehension in Real-World Applications (2021.acl-short)

Copied to clipboard

Challenge: In order to comprehensively verify the robustness and generalization of MRC models, we construct a real-world Chinese dataset - DuReader_robust .
Approach: They introduce a real-world Chinese dataset to evaluate the robustness and generalization of MRC models from three aspects: over-sensitivity, over-stability and generalisation.
Outcome: The proposed model fails to perform well on the challenge test set and may provide suggestions for future model development.
Exploring the Potential of Large Language Models in Computational Argumentation (2024.acl-long)

Copied to clipboard

Challenge: Argumentation is an essential tool in various domains, including law, public policy, and artificial intelligence.
Approach: They propose to evaluate LLMs on various computational argumentation tasks . they organize existing tasks into six main categories and standardize the format of 14 datasets .
Outcome: The proposed model performs well on argument mining and argument generation tasks.
Unsupervised Cross-Domain Prerequisite Chain Learning using Variational Graph Autoencoders (2021.acl-short)

Copied to clipboard

Challenge: Existing methods to learn prerequisite relations between concepts require annotated concept pairs during training.
Approach: They propose to use an optimized variational graph autoencoder to learn prerequisite chains in unsupervised manner using an information-rich domain and an information poor domain.
Outcome: The proposed model learns to transfer concept prerequisite relations from an information-rich domain (source domain) to an information poor domain (target domain) the annotated data and resources as well as the code will be made publicly available.
Ethos: Rectifying Language Models in Orthogonal Parameter Space (2024.findings-naacl)

Copied to clipboard

Challenge: Language models (LMs) generate toxic, biased content and reveal private training records.
Approach: They propose an efficient approach that rectifies LMs to mitigate toxicity and bias . Ethos distinguishes general beneficial and undesired knowledge when reconstructing task vectors .
Outcome: The proposed approach mitigates toxicity and bias in outputs and avoids privacy leakage.
A Query-Parallel Machine Reading Comprehension Framework for Low-resource NER (2023.findings-emnlp)

Copied to clipboard

Challenge: Named entity recognition (NER) is a fundamental task in natural language processing.
Approach: They propose a query-parallel MRC-based approach to named entity recognition . the model is trained with parameter-efficient tuning technique, making it more data-efficient .
Outcome: The proposed model performs competitively against strong baseline methods in resource-rich settings and achieves state-of-the-art results in low-resource settings.
Theory-Grounded Computational Text Analysis (2023.acl-short)

Copied to clipboard

Challenge: A broad space separates its two constituent disciplines—natural language processing and social science—which has to date been sidestepped rather than filled by applying increasingly complex computational models to problems in social science research.
Approach: They argue that computational text analysis lacks organizing principles and requires organizing methods to solve problems.
Outcome: The proposed approach is based on a review of 60 papers on computational text analysis.
MultiCite: Modeling realistic citations requires moving beyond the single-sentence single-label setting (2022.naacl-main)

Copied to clipboard

Challenge: Citation context analysis (CCA) is an important task in natural language processing that studies how and why scholars discuss each other’s work.
Approach: They propose to use a dataset of 12.6K citation contexts from 1.2K computational linguistics papers to model three important CCA phenomena.
Outcome: The proposed dataset contains 12.6K citation contexts from 1.2K computational linguistics papers and can model these phenomena.
A Self-verified Method for Exploring Simile Knowledge from Pre-trained Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Pre-trained language models (PLMs) have succeeded in natural language processing because they learn generic knowledge from a large corpus.
Approach: They propose a method that allows pre-trained language models to explore simile knowledge from PLMs . they enhance PLM models with a multi-level simile recognition task that evaluates similes aplenty .
Outcome: The proposed method can explore more accurate simile knowledge for PLMs.
Better Explain Transformers by Illuminating Important Information (2024.findings-eacl)

Copied to clipboard

Challenge: Existing explanations focus on the input and output of the Transformers, resulting in confusing results.
Approach: They propose to highlight important information and eliminate irrelevant information by a refined information flow on top of the layer-wise relevance propagation method.
Outcome: The proposed method outperforms baseline models on classification and question-answering datasets with over 3% to 33% improvement on explanation metrics.
Dynamic Routing Transformer Network for Multimodal Sarcasm Detection (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for multimodal sarcasm detection rely on fixed architectures to capture cross-modal incongruity.
Approach: They propose a method that uses dynamic paths to activate different routing transformer modules with hierarchical co-attention adapting to cross-modal incongruity.
Outcome: The proposed method is compared to state-of-the-art methods on a public dataset.
MultiMatch: Multihead Consistency Regularization Matching for Semi-Supervised Text Classification (2025.emnlp-main)

Copied to clipboard

Challenge: **MultiMatch** is a semi-supervised learning (SSL) algorithm that combines co-training and consistency regularization with pseudo-labeling.
Approach: They propose a semi-supervised learning algorithm that integrates co-training and consistency regularization with pseudo-labeling.
Outcome: The proposed algorithm outperforms the second-best approach on 8 out of 10 setups from 5 natural language processing datasets and outperformed the second best by 3.26%.
CMU-MOSEAS: A Multimodal Language Dataset for Spanish, Portuguese, German and French (2020.emnlp-main)

Copied to clipboard

Challenge: Existing datasets in multimodal language are limited and disproportionately affect native speakers of other languages . authors propose a large-scale dataset for Spanish, Portuguese, German and French .
Approach: They propose a large-scale multimodal language dataset for Spanish, Portuguese, German and French.
Outcome: The proposed dataset is the largest of its kind with 40,000 total labelled sentences . it covers a diverse set topics and speakers and carries supervision of 20 labels including sentiment, emotions, and attributes.
A Unified Span-Based Approach for Opinion Mining with Syntactic Constituents (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for fine-grained opinion mining (OM) are based on span-based annotations, but they are not effective.
Approach: They propose a unified span-based approach for the end-to-end OM setting using syntactic constituents and multi-task learning to integrate them into the proposed model.
Outcome: The proposed approach achieves significant improvements over previous work on the MPQA 2.0 dataset and reduces the number of wrongly-predicted opinion expressions and roles.
PM2F2N: Patient Multi-view Multi-modal Feature Fusion Networks for Clinical Outcome Prediction (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods focused on time series data but ignored clinical notes . fusion of multi-modal features of patients from different views is not feasible due to the time series and clinical notes data being stored as time series.
Approach: They propose to combine time series and clinical notes to fuse multi-modal features of patients from different perspectives using graph neural networks.
Outcome: The proposed method is superior to existing models on MIMIC-III benchmark.
NarrowBERT: Accelerating Masked Language Model Pretraining and Inference (2023.acl-short)

Copied to clipboard

Challenge: Large-scale language model pretraining is expensive as the models and pretraining corpora have become larger over time.
Approach: They propose a modified transformer encoder that increases throughput for masked language model pretraining by more than 2x.
Outcome: The proposed model increases throughput on IMDB and Amazon reviews classification and CoNLL NER tasks by 3.5x with minimal performance degradation.
End-to-End Neural Word Alignment Outperforms GIZA++ (2020.acl-main)

Copied to clipboard

Challenge: Word alignment was once a core unsupervised learning task in natural language processing . but word alignment still plays an important role in interactive applications of neural machine translation, such as annotation transfer and lexicon injection.
Approach: They propose to use a Transformer model to train an unsupervised word alignment model.
Outcome: The proposed method outperforms GIZA++ on three data sets and is tightly integrated and does not affect translation quality.
TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification (2020.findings-emnlp)

Copied to clipboard

Challenge: Modern NLP systems are typically ill-equipped when applied to noisy user-generated text.
Approach: They propose a new evaluation framework consisting of seven Twitter-specific classification tasks.
Outcome: The proposed framework is based on seven heterogeneous Twitter-specific classification tasks.
BehancePR: A Punctuation Restoration Dataset for Livestreaming Video Transcript (2022.findings-naacl)

Copied to clipboard

Challenge: a growing number of livestreaming videos provide useful knowledge with exceptional visual demonstrations.
Approach: They propose a human-annotated corpus for punctuation restoration in livestreaming video transcripts . they show popular natural language processing tools underperform on sentence boundary detection .
Outcome: The proposed dataset shows that natural language processing tools underperform on sentence boundary detection on livestreaming video transcripts.
Step by Step Loss Goes Very Far: Multi-Step Quantization for Adversarial Text Attacks (2023.eacl-main)

Copied to clipboard

Challenge: Existing gradient-based attacks quantize all tokens in a text at once, which creates a significant gap between adversarial loss for continuous and discrete text representations.
Approach: They propose a gradient-based attack that quantizes tokens one by one and reoptimizes adversarial example after each quantization.
Outcome: The proposed method outperforms other approaches on various natural language processing tasks.
VISPool: Enhancing Transformer Encoders with Vector Visibility Graph Neural Networks (2024.findings-acl)

Copied to clipboard

Challenge: Existing graph-based graph construction methods rely on static graphs and are not scalable with increasing document and word counts.
Approach: They propose a dynamic graph construction method based on vector visibility graphs (VVGs) they propose scalable model architecture that integrates VVG convolutional networks into transformer pipelines.
Outcome: The proposed model outperforms baseline models on the GLUE benchmark datasets.
Surveying the Dead Minds: Historical-Psychological Text Analysis with Contextualized Construct Representation (CCR) for Classical Chinese (2024.emnlp-main)

Copied to clipboard

Challenge: Humans have produced written language for thousands of years, but most computational work is focused on contemporary languages and corpora.
Approach: They propose a pipeline for historical-psychological text analysis in classical Chinese . they propose an indirect contrastive learning approach that fine-tunes pre-trained models .
Outcome: The proposed pipeline outperforms word-embedding-based approaches across all tasks and exceeds prompting with GPT-4 in most tasks.
Reconstruction Attack on Instance Encoding for Language Understanding (2021.emnlp-main)

Copied to clipboard

Challenge: Existing private learning schemes which protect data privacy can be used to train models using instance encoding.
Approach: They propose to recover the private training data and use it to break a private learning scheme TextHide.
Outcome: The proposed attack would advance privacy-preserving machine learning in the context of natural language processing.
Sentence-Level Resampling for Named Entity Recognition (2022.naacl-main)

Copied to clipboard

Challenge: named entity recognition (NER) tasks are often dominated by the majority of non-entity tokens in text . a data imbalance problem is causing the NER models to ignore named entities .
Approach: They propose a set of sentence-level resampling methods to reduce data imbalance . they use a training sentence to compute the importance of each training sentence based on its tokens and entities .
Outcome: The proposed methods outperform sub-sentence-level resampling, data augmentation, and loss functions on multiple corpora.
Active2 Learning: Actively reducing redundancies in Active Learning methods for Sequence Tagging and Machine Translation (2021.naacl-main)

Copied to clipboard

Challenge: Existing approaches to deep learning for NLP require large amounts of labeled data.
Approach: They propose an approach that iteratively selects a small number of examples for expert annotation based on their estimated utility in training the model.
Outcome: The proposed approach reduces the data requirements of state-of-the-art AL strategies by 3-25% on multiple NLP tasks while achieving the same performance with virtually no additional computation overhead.
Us vs. Them: A Dataset of Populist Attitudes, News Bias and Emotions (2021.eacl-main)

Copied to clipboard

Challenge: Populist rhetoric has risen across the political sphere in recent years, but computational approaches to it have been scarce.
Approach: They propose a dataset of 6861 reddit comments annotated for populist attitudes and a set of multi-task learning models that leverage emotion and group identification as auxiliary tasks.
Outcome: The proposed models leverage emotion and group identification as auxiliary tasks to model populist rhetoric tasks.
SYNTHVERIFY: Enhancing Zero-Shot Claim Verification through Step-by-Step Synthetic Data Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for claim verification are inefficient or rely on external documents.
Approach: They propose a step-by-step prompting-based synthetic data generation framework to enhance zero-shot claim verification.
Outcome: The proposed framework bridges LLMs’ knowledge gaps in specialized domains without access to external corpora or sacrificing generalizability.
An Evaluation of Progressive Neural Networksfor Transfer Learning in Natural Language Processing (2020.lrec-1)

Copied to clipboard

Challenge: Fine-tuning suffers from catastrophic forgetting, a problem exacerbated in natural language processing (NLP).
Approach: They propose to use progressive neural networks to re-use previously learned knowledge when learning new tasks.
Outcome: The proposed approach improves on common NLP tasks across a range of architectures, datasets, and tasks.
Streaming word similarity mining on the cheap (D18-1)

Copied to clipboard

Challenge: Existing methods to estimate word similarities are to embed words in vector space and then calculate similarities between corresponding vectors.
Approach: They propose a method that explicitly counts second-order co-occurrences to estimate word similarities from streams.
Outcome: The proposed method is scalable, converges rapidly, behaves robustly under parameter changes, and captures word similarities on par with state-of-the-art word embeddings.
A Multi-Format Transfer Learning Model for Event Argument Extraction via Variational Information Bottleneck (2022.coling-1)

Copied to clipboard

Challenge: Event argument extraction (EAE) aims to extract arguments with given roles from texts.
Approach: They propose a multi-format transfer learning model with variational information bottleneck to learn from existing datasets.
Outcome: The proposed model improves on three benchmark datasets and obtains state-of-the-art performance on EAE.
Towards Understanding Gender-Seniority Compound Bias in Natural Language Generation (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies have not investigated how gender biases in natural language processing (NLP) are compounded with other societal biase.
Approach: They propose a framework for probing compound bias by examining seniority in pre-trained neural generation models.
Outcome: The proposed framework amplifies bias by considering women as junior and men as senior more often than ground truth in both domains.
Investigating the Working of Text Classifiers (C18-1)

Copied to clipboard

Challenge: Text classification is one of the most widely studied tasks in natural language processing.
Approach: They propose to use large multilayer neural network models to compose meaning of sentences . they propose to disincentivize focusing on key lexicons to improve classification accuracy .
Outcome: The proposed models learn to compose the meaning of the sentences or focus on key lexicons for classifying the document.
A Review on Deep Learning Techniques Applied to Answer Selection (C18-1)

Copied to clipboard

Challenge: Existing deep learning methods for answer selection are not feature engineering or expensive external resources.
Approach: They propose to use deep learning methods to analyze and predict answer quality . they use a set of candidate answers to identify which of the candidates answers the question correctly.
Outcome: The proposed methods produce impressive performance without feature engineering or expensive external resources.
Collaboration or Corporate Capture? Quantifying NLP’s Reliance on Industry Artifacts and Contributions (2024.acl-long)

Copied to clipboard

Challenge: EMNLP 2022 citations are three times greater than expected for pre-trained models . industry participation in the Association of Computational Linguistics (ACL) anthology has increased 180% from 2017 to 2022.
Approach: They surveyed 100 papers published at EMNLP 2022 to determine the ratio of their citations to industry models.
Outcome: a new study shows that industry citations are three times greater than expected . the study aims to better understand whether industry collaboration is still collaboration . industry participation in the 2023 AI index report is the top takeaway .
Calibrating Structured Output Predictors for Natural Language Processing (2020.acl-main)

Copied to clipboard

Challenge: Several modern machine-learning based NLP systems can provide a confidence score with their output predictions.
Approach: They propose a general calibration scheme for output entities of interest in NLP applications that can be used to calibrate confidence scores.
Outcome: The proposed calibration scheme outperforms current calibration techniques for Named Entity Recognition, Part-of-speech tagging and Question Answering systems.
Show Your Work with Confidence: Confidence Bands for Tuning Curves (2024.naacl-long)

Copied to clipboard

Challenge: a rush to scale up has left us with large, costly language models and little understanding of how different designs compare.
Approach: They propose a method to construct valid confidence bands for tuning curves . they validated their method with ablations and analyze the effect of sample size .
Outcome: The proposed method shows that bootstrap confidence bands do not approximate their target confidence.
Partially-Random Initialization: A Smoking Gun for Binarization Hypothesis of BERT (2022.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained BERT has been used for natural language processing tasks but its performance is limited by memory and computational complexity.
Approach: They propose to use pre-trained BERT to achieve decent accuracy . they propose to combine binary BERT with a randomly-initialized encoder .
Outcome: The proposed model achieves state-of-the-art on GLUE and SQuAD benchmarks.
FLiText: A Faster and Lighter Semi-Supervised Text Classification with Convolution Networks (2021.emnlp-main)

Copied to clipboard

Challenge: obtaining large amounts of labeled data is expensive.
Approach: They develop a semi-supervised learning framework called FLiText which improves text classification accuracy.
Outcome: The proposed framework improves accuracy of lightweight models on IMDb, Yelp-5, and Yahoo! Answer . the framework improve accuracy by 6.59%, 3.94%, and 3.22% on the datasets of IMDa, Yep-5 and Yahoo. Answer compared with the fully supervised method on the full dataset .
R2F: A General Retrieval, Reading and Fusion Framework for Document-level Natural Language Inference (2022.emnlp-main)

Copied to clipboard

Challenge: Document-level natural language inference (DOCNLI) is a new task in natural language processing.
Approach: They propose a document-level natural language inference framework that fuses sentence-level tasks into a set of sentence-based tasks.
Outcome: The proposed framework improves interpretability and performance with evidence.
Beyond Canonical Fine-tuning: Leveraging Hybrid Multi-Layer Pooled Representations of BERT for Automated Essay Scoring (2024.lrec-main)

Copied to clipboard

Challenge: Existing work on automated essay scoring focuses on capturing deep semantic features but are limited to lower-level textual features.
Approach: They propose to use BERT's multi-layer architecture to leverage hierarchical linguistic information from its intermediate layers to improve overall essay scoring performance.
Outcome: The proposed model outperforms the standard model with the default output on the ASAP AES dataset.
Neural Keyphrase Generation via Reinforcement Learning with Adaptive Rewards (P19-1)

Copied to clipboard

Challenge: Existing generative models generate too few keyphrases, but they often generate too many . et al. (2017) propose a reinforcement learning approach for keyphrase generation .
Approach: They propose a reinforcement learning approach that encourages a model to generate sufficient keyphrases with an adaptive reward function.
Outcome: The proposed method improves state-of-the-art generative models with conventional and new evaluation methods on real-world datasets.
Cross-Domain Classification of Moral Values (2022.findings-naacl)

Copied to clipboard

Challenge: Existing methods to identify moral values in text can be challenging for transferring knowledge between domains.
Approach: They compare a deep learning model with a domain-specific value classifier to find out whether it can transfer knowledge to new domains.
Outcome: The proposed model can generalize and transfer knowledge to novel domains, but introduce catastrophic forgetting.
A Learning-Exploring Method to Generate Diverse Paraphrases with Multi-Objective Deep Reinforcement Learning (2020.coling-main)

Copied to clipboard

Challenge: Paraphrase generation is of great importance for many downstream tasks in natural language processing.
Approach: They propose a method to generate sentences as learning objectives from the learned data distribution and employ reinforcement learning to combine these new learning objectives for model training.
Outcome: The proposed method gains significant diversity and improves generation quality over state-of-the-art datasets.
Learning Context-Sensitive Convolutional Filters for Text Processing (D18-1)

Copied to clipboard

Challenge: Convolutional neural networks (CNNs) are a popular building block for natural language processing . despite their success, most existing CNN models share the same learned set of filters for all input sentences.
Approach: They propose to use a meta network to learn context-sensitive convolutional filters for text processing by using a bidirectional filter generation mechanism.
Outcome: The proposed framework outperforms standard and attention-based CNN models on four different tasks.
Benchmarking Language Models for Code Syntax Understanding (2022.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained language models capture the syntactic rules of natural languages without fine-tuning on syntax understanding tasks.
Approach: They propose a benchmarking test to compare pre-trained language models with a large-scale dataset of programs annotated with syntactic relationships in their corresponding abstract syntax trees.
Outcome: The proposed model fails to match baselines based on positional offsets and keywords.
INSET: Sentence Infilling with INter-SEntential Transformer (2020.acl-main)

Copied to clipboard

Challenge: Missing sentence generation fosters a wide range of applications in natural language generation . Developing models for sentence infilling can potentially facilitate many text generation applications .
Approach: They propose a framework to decouple the problem from natural language processing . they propose generating missing sentences that can syntactically and semantically bridge context .
Outcome: The proposed model learns a sentence representation and generates 'missing sentences' the proposed model can be used for document auto-completion and meeting note expansion .
Automatic Text Simplification for Social Good: Progress and Challenges (2021.findings-acl)

Copied to clipboard

Challenge: ATS has been promoted as a natural language processing task since the 1990s . but since 2010, the field has been focusing on building complex end-to-end neural architectures based on ATS .
Approach: They propose to use automated text simplification (ATS) to make texts more accessible to people with disabilities . they argue that lack of high-quality TS datasets and standardized evaluation procedures are barriers .
Outcome: The proposed neural ATS systems are based on a new set of TS datasets and a standardized evaluation procedure.
Effective Self-Mining of In-Context Examples for Unsupervised Machine Translation with LLMs (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive performance on a wide range of natural language processing tasks.
Approach: They propose an unsupervised approach to mine in-context examples for machine translation (MT) they use word-level mining to acquire word translations that are then used to perform sentence-level mines .
Outcome: The proposed approach outperforms state-of-the-art methods on 288 directions on 287 languages and is based on word-level mining and sentence-level extraction.
BanLemma: A Word Formation Dependent Rule and Dictionary Based Bangla Lemmatizer (2023.findings-emnlp)

Copied to clipboard

Challenge: Lemmatization holds significance in both natural language processing (NLP) and linguistics due to the highly inflected nature and morphological richness of Bangla text.
Approach: They propose linguistic rules for lemmatization and utilize a dictionary along with the rules to design a lemma specifically for Bangla.
Outcome: The proposed system achieves 96.36% accuracy when tested against a manually annotated test dataset.
Towards Efficient NLP: A Standard Evaluation and A Strong Baseline (2022.naacl-main)

Copied to clipboard

Challenge: Rather than pursuing the reachless SOTA accuracy, researchers are focusing on model efficiency and usability.
Approach: They propose an evaluation and a public leaderboard for efficient NLP models that depicts the Pareto Frontier for various language understanding tasks.
Outcome: The proposed model outperforms or performs on par with SOTA compressed and early exiting models.
BERT-EMD: Many-to-Many Layer Mapping for BERT Compression with Earth Mover’s Distance (2020.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models have been proposed and applied to many NLP tasks, yielding state-of-the-art performance, but high storage and computational costs obstruct them to be effectively deployed on resource-constrained devices and real-time applications.
Approach: They propose a BERT distillation method which allows each intermediate student layer to learn from any intermediate teacher layers.
Outcome: The proposed method can learn from different teacher layers adaptively for different NLP tasks.
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models that assess explicit and implicit biases are based on a single scenario . a dataset of 450 offensive progressions contains 2,700 sentences of varying severity .
Approach: They evaluate a dataset of offensive progressions that contain 2,700 sentences . they find that even the best-performing models detect bias inconsistently .
Outcome: The proposed dataset shows that even the best-performing models detect bias inconsistently . aligning models with human judgments on STOP can improve answer rates on sensitive tasks by 191% .
HacRED: A Large-Scale Relation Extraction Dataset Toward Hard Cases in Practical Applications (2021.findings-acl)

Copied to clipboard

Challenge: Relation extraction (RE) is an essential topic in natural language processing and has attracted extensive attention.
Approach: They propose a case-oriented construction framework to build a hard case relation extraction dataset with 65,225 relational facts annotated from 9,231 documents.
Outcome: The proposed model achieves a high 96% F1 score on data quality and is far lower than humans.
How Well Do LLMs Handle Cantonese? Benchmarking Cantonese Capabilities of Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Cantonese has scant representation in NLP research, especially compared to other languages from similarly developed regions.
Approach: They propose to evaluate Cantonese LLM performance in factual generation, mathematical logic, complex reasoning, and general knowledge in Cantonesian.
Outcome: The proposed models will evaluate Cantonese's performance in factual generation, mathematical logic, complex reasoning, and general knowledge in Cantone.
Enhancing Language Representation with Constructional Information for Natural Language Understanding (2023.acl-long)

Copied to clipboard

Challenge: Recent advances in natural language processing focus on acquiring lexico-semantic information.
Approach: They propose a construction grammar which highlights the pairings of form and meaning to enrich language representation.
Outcome: The proposed model is superior to existing models on a variety of NLU tasks.
2INER: Instructive and In-Context Learning on Few-Shot Named Entity Recognition (2023.findings-emnlp)

Copied to clipboard

Challenge: Named Entity Recognition (NER) tasks are a fundamental task of natural language processing (NLP).
Approach: They propose a text-to-text framework for Few-Shot Named Entity Recognition (NER) that employs instruction finetuning and auxiliary tasks to enhance the model's understanding of entity types in the overall semantic context of a sentence.
Outcome: The proposed framework outperforms existing Few-Shot NER methods and remains competitive with state-of-the-art NER algorithms.
WavLLM: Towards Robust and Adaptive Speech Large Language Model (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have expanded their scope to encompass multimodal functions.
Approach: They propose a robust and adaptive speech large language model with dual encoders . they validate the model on universal speech benchmarks and apply it to specialized speech-question-answer datasets based on a CoT approach .
Outcome: The proposed model achieves state-of-the-art performance across a range of speech tasks on the same model size.
Benchmarking Intersectional Biases in NLP (2022.naacl-main)

Copied to clipboard

Challenge: Recent work on fairness of machine learning models has focused on how to debias, but research on the fairness and performance of biased/debiased models on downstream prediction tasks has been limited.
Approach: They assess intersectional bias - fairness across multiple demographic dimensions . they highlight possible causes and make recommendations for future NLP debiasing research.
Outcome: The proposed approaches fare well in terms of fairness-accuracy trade-off, but are unable to effectively alleviate bias in downstream tasks.
Mitigating Gender Bias Amplification in Distribution by Posterior Regularization (2020.acl-main)

Copied to clipboard

Challenge: Recent studies show that data-driven machine learning models carry societal biases in the dataset they trained on.
Approach: They propose to calibrate top predictions of a model by injecting corpus-level constraints to ensure that the gender disparity is not amplified.
Outcome: The proposed method can almost remove bias amplification in the distribution with little loss of performance.
HILL: Hierarchy-aware Information Lossless Contrastive Learning for Hierarchical Text Classification (2024.naacl-long)

Copied to clipboard

Challenge: Existing self-supervised methods in natural language processing rely on augmentation rules to generate contrastive samples.
Approach: They propose a hierarchy-aware information lossless contrastive learning scheme that uses syntactic information reserved in the input sample and fused during the learning process.
Outcome: The proposed learning scheme is superior to existing methods in hierarchical text classification . the proposed learning system is based on a structure encoder and a text encoder .
Learning Physical Common Sense as Knowledge Graph Completion via BERT Data Augmentation and Constrained Tucker Factorization (2020.emnlp-main)

Copied to clipboard

Challenge: Physical commonsense learning is an essential part of human-robot interaction . existing methods of learning physical commons sense suffer from generalization .
Approach: They propose to use physical commonsense learning as a knowledge graph completion problem to better use latent relationships among training samples.
Outcome: The proposed method outperforms existing methods in the human-robot interaction problem.
Bridging the Language Gaps in Large Language Models with Inference-Time Cross-Lingual Intervention (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to address performance gaps in LLMs rely on pretraining or fine-tuning, which are resource-intensive.
Approach: They propose a framework that aligns LLMs' internal representations with those of high-performing languages during inference.
Outcome: The proposed framework improves performance on low-performing (source) languages by aligning their internal representations with those of high-performing languages during inference.
Human and LLM-Based Resume Matching: An Observational Study (2025.findings-naacl)

Copied to clipboard

Challenge: Resume matching assesses the extent to which candidates qualify for jobs based on the content of resumes.
Approach: They compare GPT-4 and human ratings for resumes submitted to job openings from diverse fields using real-world evaluation criteria.
Outcome: The proposed model improves the quality of LLM ratings and does not show bias.
ChatEL: Entity Linking with Chatbots (2024.lrec-main)

Copied to clipboard

Challenge: Entity Linking (EL) is a challenging task in natural language processing . existing approaches focus on creating elaborate contextual models that are unwieldy and difficult to train .
Approach: They propose a framework to prompt LLMs to return accurate results for Entity Linking . they use a three-step framework to generate a set of EL models that can be open-source .
Outcome: The proposed framework improves the average F1 performance across 10 datasets by more than 2%.
Estimating Agreement by Chance for Sequence Annotation (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on chance correction for sequence annotation tasks lack a chance corrected agreement metric.
Approach: They propose a model for generating random annotations which serves as the foundation for estimating chance agreement in sequence annotation tasks.
Outcome: The proposed model is validated in simulation and corpus-based evaluation.
HG2Vec: Improved Word Embeddings from Dictionary and Thesaurus Based Heterogeneous Graph (2022.coling-1)

Copied to clipboard

Challenge: Existing models that learn word embeddings rely on a large corpus of data . however, these models require massive time and space for data pre-processing and training .
Approach: They propose a model that learns word embeddings utilizing only dictionaries and thesauri . they exploit a new context-focused loss model that models transitive relationships between word pairs .
Outcome: The proposed model reaches the state-of-art on multiple word similarity and relatedness benchmarks.
Learning from Impairment: Leveraging Insights from Clinical Linguistics in Language Modelling Research (2025.coling-main)

Copied to clipboard

Challenge: Using neurolinguistics and aphasiology, we examine the theoretical underpinnings of some influential linguistically motivated training approaches targeting the syntactic domain.
Approach: They examine the theoretical underpinnings of linguistically motivated training approaches derived from neurolinguistics and aphasiology to develop human-like learning strategies for language models.
Outcome: The proposed frameworks can be used to improve the recovery and generalization of linguistic skills in aphasia treatment and to develop human-like learning strategies.
Is the Lottery Fair? Evaluating Winning Tickets Across Demographics (2021.findings-acl)

Copied to clipboard

Challenge: Recent studies suggest weight pruning compromises fairness of machine learning models . however, no empirical evaluation has been done in the context of natural language processing.
Approach: They evaluate the fairness of lottery ticket extraction through layer-wise and global weight pruning across three languages and two tasks.
Outcome: The proposed model is compared with two text classification datasets annotated with demographic information.
Frugal Prompting for Dialog Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are used in natural language processing tasks with an unrealistic speed and effectiveness.
Approach: They propose more compact ways of providing dialog history information while ensuring good performance and reducing model’s inference-API costs.
Outcome: The proposed models have the optimal usable-information density while maintaining good performance and reducing model’s inference-API costs.
KORE 50ˆDYWC: An Evaluation Data Set for Entity Linking Based on DBpedia, YAGO, Wikidata, and Crunchbase (2020.lrec-1)

Copied to clipboard

Challenge: A major domain of research in natural language processing is named entity recognition and disambiguation (NERD).
Approach: They extend a widely-used data set to include NERD tasks for DBpedia and YAGO, Wikidata and Crunchbase.
Outcome: The extended data set allows for a broader spectrum of evaluation.
Word-Level Loss Extensions for Neural Temporal Relation Classification (C18-1)

Copied to clipboard

Challenge: Unsupervised pre-trained word embeddings are used for many tasks in natural language processing to leverage unlabeled textual data.
Approach: They extend the model's task loss with an unsupervised auxiliary loss on the word-embedding level of the model to ensure that the learned word representations contain both task-specific features and more general features.
Outcome: The proposed model improves on the task of extracting narrative containment relations from clinical records using a general-domain part-of-speech tagger as linguistic resource.
On the Impact of Temporal Concept Drift on Model Explanations (2022.findings-emnlp)

Copied to clipboard

Challenge: Explanation faithfulness of model predictions is typically evaluated on held-out data from the same temporal distribution as the training data.
Approach: They examine the impact of temporal variation on model explanations extracted by eight feature attribution methods and three select-then-predict models across six text classification tasks.
Outcome: The proposed method shows the most robust faithfulness scores across datasets and in asynchronous settings.
Building a Macro Chinese Discourse Treebank (L18-1)

Copied to clipboard

Challenge: Discourse structure analysis is an important research topic in natural language processing.
Approach: They propose to construct a macro discourse structure framework and annotate 147 Newswire articles.
Outcome: The proposed framework can lay the foundation for further analysis of macro discourse structure.
Text2Tabular – Reconstructing Tabular Research Data from Scientific Publications (2026.acl-long)

Copied to clipboard

Challenge: Text2Tabular reconstructs research datasets from scientific literature using advanced natural language processing and statistical modeling.
Approach: Text2Tabular reconstructs research datasets from scientific publications using natural language processing and statistical modeling.
Outcome: Text2Tabular reconstructs scientific literature-based datasets using natural language processing and statistical modeling.
It’s All in the Heads: Using Attention Heads as a Baseline for Cross-Lingual Transfer in Commonsense Reasoning (2021.findings-acl)

Copied to clipboard

Challenge: gilbert et al.: commonsense reasoning is a key problem in natural language processing but its capabilities are still unstudied. gilland eetal.: a new approach to commonsensible reasoning is needed to solve the problem.
Approach: They propose a method which trains a linear classifier with weights of multi-head attention as features and a multilingual Winograd Schema corpus to measure cross-lingual generalization ability.
Outcome: The proposed approach performs competitively with recent approaches even when applied to other languages in a zero-shot manner.
Understanding Attention for Text Classification (2020.acl-main)

Copied to clipboard

Challenge: Existing studies have focused on whether local attention weights reflect the importance of input representations.
Approach: They propose to analyze for each word token the following two quantities: its polarity score and its attention score, where the latter is a global assessment on the token’s significance.
Outcome: The proposed model can be improved under conditions where the interplay between the two quantities can contribute towards model performance.
Exploring Diverse Expressions for Paraphrase Generation (D19-1)

Copied to clipboard

Challenge: Existing neural paraphrase generation methods focus on single paraphrases while ignoring the fact that diversity is essential for enhancing generalization capability and robustness of downstream applications.
Approach: They propose a novel approach with two discriminators and multiple generators to generate a variety of different paraphrases.
Outcome: The proposed model gains significant diversity and improves quality over state-of-the-art datasets.
ProQA: Structural Prompt-based Pre-training for Unified Question Answering (2022.naacl-main)

Copied to clipboard

Challenge: Existing QA research on question answering is focused on specific question types, knowledge domains, or reasoning skills.
Approach: They propose a unified QA paradigm that solves various tasks through a single model.
Outcome: The proposed model improves QA-centric ability on 11 QA benchmarks.
Slender-Mamba: Fully Quantized Mamba in 1.58 Bits From Head to Toe (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) have achieved significant performance improvements in natural language processing domain, but require large computational resources for training and inference.
Approach: They propose to use a language model architecture based on State-Space Models to quantify embedding and projection layers of a model with 150 B tokens from scratch.
Outcome: The proposed language model architecture reduces costs by compressing context windows during inference while reducing the cost of training and inference.
KoBEST: Korean Balanced Evaluation of Significant Tasks (2022.coling-1)

Copied to clipboard

Challenge: a well-formulated benchmark allows objective and precise evaluation of diverse models.
Approach: They propose a benchmark for Korean balanced evaluation of significant tasks that requires advanced Korean linguistic knowledge.
Outcome: The proposed benchmarks are based on five Korean-language downstream tasks . the data is annotated by humans and thoroughly reviewed to guarantee high data quality.
Modeling the Unigram Distribution (2021.findings-acl)

Copied to clipboard

Challenge: a novel model for estimating the unigram distribution is proposed for a language . the model is based on the word type and word type distributions, but does not consider contextual information.
Approach: They propose a neuralization of Goldwater's (2011) model for estimating the unigram distribution in a language.
Outcome: The proposed model outperforms character-level models in estimating the unigram distribution in a language.
Large Language Models for Anomaly and Out-of-Distribution Detection: A Survey (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated their effectiveness in natural language processing but also in broader applications due to their advanced comprehension and generative capabilities.
Approach: They propose a taxonomy to categorize existing approaches into two classes based on the role played by LLMs.
Outcome: The proposed taxonomy categorizes existing approaches into two classes based on the role played by LLMs.
Differential Privacy for Text Analytics via Natural Text Sanitization (2021.findings-acl)

Copied to clipboard

Challenge: Existing text sanitization mechanisms provide low utility, as cursed by the high-dimensional text representation.
Approach: They propose to use sanitized texts to samaritize training data . they propose to retrain and fine-tune the senitization-aware language model .
Outcome: The proposed approach enables privacypreserving natural language processing over the BERT language model with promising utility.
Benchmarking Large Language Models on CFLUE - A Chinese Financial Language Understanding Evaluation Dataset (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have revolutionized natural language processing (NLP) there is an urgent need for new benchmarks to keep pace with the development of LLMs.
Approach: They propose a benchmark to assess the capability of large language models (LLMs) they use a dataset to provide both knowledge assessment and application assessment .
Outcome: The proposed benchmark provides datasets tailored for knowledge assessment and application assessment.
MultiMWE: Building a Multi-lingual Multi-Word Expression (MWE) Parallel Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Existing bilingual or multi-lingual MWE corpora are limited for multilingual use . only 871 pairs of English-German MWEs are available for research .
Approach: They present a collection of bilingual and multi-lingual MWEs extracted from parallel corpora.
Outcome: The available bilingual or multi-lingual MWE corpus is very limited . the collection is a small collection of 871 pairs of English-German MWEs .
STAPI: An Automatic Scraper for Extracting Iterative Title-Text Structure from Web Documents (2022.lrec-1)

Copied to clipboard

Challenge: Formal documents are organized into sections of text, each with a title . but there is no corpus of web documents annotated with titles and prose texts . cnn.com's john mccarthy and daniel mclears are working on a new title-text dataset .
Approach: They propose a first title-text dataset on web documents that incorporates a wide variety of domains to facilitate downstream training.
Outcome: The proposed system outperforms baseline models in terms of title-text identification.
Cross-type French Multiword Expression Identification with Pre-trained Masked Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Multiword expressions (MWEs) have linguistic features that distinguish them from regular word groupings.
Approach: They propose a combination of two systems that learn verbal multiword expressions and non-verbal MWEs to improve performance on a cross-type dataset .
Outcome: The proposed system improves the F1 score on a french treebank with VMWEs and nVMWES training data.
Transductive Learning of Neural Language Models for Syntactic and Semantic Analysis (D19-1)

Copied to clipboard

Challenge: despite its practical advantages, transductive learning is underexplored in natural language processing . despite the simplicity of the technique, it is understudied in natural languages .
Approach: They conduct an empirical study of transductive learning for neural models . they fine-tune language models on an unlabeled test set to obtain test-set-specific word representations.
Outcome: The proposed method improves state-of-the-art neural models in syntactic and semantic tasks.
ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences (2024.acl-long)

Copied to clipboard

Challenge: Current large language models (LLMs) are ineffective in learning domain knowledge and aligning with human preference.
Approach: They propose a benchmark LLM for Chinese medical domain that uses pre-training, supervised fine-tuning and RLHF to train LLMs.
Outcome: The proposed LLM performs better than existing LLMs in the Chinese medical domain.
Towards a Linked Open Data Edition of Sumerian Corpora (L18-1)

Copied to clipboard

Challenge: Linguistic Linked Open Data (LLOD) is a flourishing line of research in the language resource community . existing LLOD standards and vocabularies are not widely used in this community despite its popularity .
Approach: They propose to use Linguistic Linked Open Data to link a Sumerian corpus with lexical resources . they use a linguistically annotated archive to create a corpus of cuneiform texts .
Outcome: The proposed LLOD framework is used in assyriology, with philological resources underrepresented . the proposed framework is based on a linguistically annotated corpus of Sumerian texts .
PTCSpell: Pre-trained Corrector Based on Character Shape and Pinyin for Chinese Spelling Correction (2023.findings-acl)

Copied to clipboard

Challenge: Chinese spelling correction (CSC) is a task which detects incorrect characters in Chinese text and corrects them.
Approach: They propose to pre-train a Chinese spelling correction corrector under the detector-corrector architecture and propose to capture pronunciation and shape information in Chinese characters.
Outcome: The proposed corrector achieves an average of 5.8% F1 improvements over state-of-the-art methods, verifying its effectiveness.
Joint Modelling of Emotion and Abusive Language Detection (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for abuse detection focus on linguistic properties of comments and online communities of users, disregarding the emotional state of the users and how this might affect their language.
Approach: They propose to combine emotion and abusive language detection to create a multi-task learning framework that allows one task to inform the other.
Outcome: The proposed model improves on the previous models, incorporating affective features into the learning framework.
Contextualized Perturbation for Textual Adversarial Attack (2021.naacl-main)

Copied to clipboard

Challenge: Existing techniques for generating adversarial examples are driven by local heuristic rules that are agnostic to the context, resulting in unnatural and ungrammatical outputs.
Approach: They propose a ContextuaLized AdversaRial Example generation model that generates fluent and grammatical outputs through a mask-then-infill procedure.
Outcome: The proposed model outperforms baseline models in terms of attack success rate, textual similarity, fluency and grammaticality.
ConvTextTM: An Explainable Convolutional Tsetlin Machine Framework for Text Classification (2022.lrec-1)

Copied to clipboard

Challenge: Recent advances in natural language processing (NLP) have reshaped the industry . complexity of such models makes them a “black box” and can cause ethical concerns .
Approach: They propose a convolutional TM architecture that breaks down text into a sequence of fragments . they propose to use a tokenization scheme to bind the tokens to the text fragments.
Outcome: The proposed architecture improves on a set of text fragments and eliminates the need for a corpus-specific vocabulary.
From Text Segmentation to Enhanced Representation Learning: A Novel Approach to Multi-Label Classification for Long Texts (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing models rely on pre-trained language models, which have a maximum input sequence length of 512 tokens, and therefore have 'input length limitation'.
Approach: They propose a text segmentation algorithm which guarantees to produce the optimal segmentation to address the issue of input length limitation caused by PLMs.
Outcome: The proposed method improves both text and label representations on MLTC datasets, unraveling the intricate correlations between texts and labels.
Consistent Accelerated Inference via Confident Adaptive Transformers (2021.emnlp-main)

Copied to clipboard

Challenge: Amortized or approximate computational methods increase efficiency, but can result in unpredictable performance costs.
Approach: They propose a method that increases computational efficiency while guaranteeing a specifiable degree of consistency with the original model with high confidence.
Outcome: The proposed method improves on four classification and regression tasks and can be used to predict the performance of the proposed model.
Neural Grammatical Error Correction with Finite State Transducers (N19-1)

Copied to clipboard

Challenge: Language model based GEC (LM-GEC) is a promising alternative to SMT and neural sequence-to-sequence models.
Approach: They propose to use finite state transducers to improve LM-GEC by rescoring with neural language models.
Outcome: The proposed model outperforms the best published results on the CoNLL-2014 test set and achieves far better relative improvements over the baselines.
MultiCMET: A Novel Chinese Benchmark for Understanding Multimodal Metaphor (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing research on multimodal metaphors does not address categorizing the source and target domains in metaphors beyond the English language.
Approach: They propose a Cascading Domain Knowledge Integration benchmark to detect metaphors by introducing domain-specific lexical features.
Outcome: The proposed dataset includes 13,820 text-image pairs of advertisements with manual annotations of the occurrence of metaphors, domain categories, and sentiments metaphors convey.
Exploring BERT’s Sensitivity to Lexical Cues using Tests from Semantic Priming (2020.findings-emnlp)

Copied to clipboard

Challenge: Using English lexical stimuli, we find that BERT models show "priming" predicting a word with greater probability when the context includes a related word versus an unrelated one.
Approach: They analyze a pre-trained BERT model with tests informed by semantic priming . they find that BERT too shows "priming" predicting a word with greater probability when context includes a related word versus an unrelated one.
Outcome: The proposed model shows a tendency to be distracted by related prime words as context becomes more informative, and lower probability of related words.
Enhancing Variational Autoencoders with Mutual Information Neural Estimation for Text Generation (D19-1)

Copied to clipboard

Challenge: Existing approaches to train variational autoencoders (VAEs) have been proposed to alleviate the posterior collapse issue in NLP tasks.
Approach: They propose to introduce a mutual information term between the input and its latent variable to regularize the objective of the VAE.
Outcome: The proposed model performs better on three benchmark datasets and is comparable to state-of-the-art models.
What Would a Teacher Do? Predicting Future Talk Moves (2021.findings-acl)

Copied to clipboard

Challenge: Recent advances in natural language processing (NLP) have the ability to transform how classroom learning takes place.
Approach: They propose a task that uses the academically productive talk framework to learn strategies that make for the best learning experience.
Outcome: The proposed task outperforms baselines on academically productive talk (FTMP) and shows that it outperformed human performance on FTMP.
SciReviewGen: A Large-scale Dataset for Automatic Literature Review Generation (2023.findings-acl)

Copied to clipboard

Challenge: Existing literature review models have addressed literature review generation, but lack of large-scale datasets has been a stumbling block.
Approach: They propose to use a large-scale dataset to evaluate automatic literature review generation models.
Outcome: The proposed model can generate summaries comparable to human-written reviews while lacking detailed information.
DET: A Dual-Encoding Transformer for Relational Graph Embedding (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to graph representation only consider the local neighbors, sacrificing the Transformer’s ability to attend to elements at any distance.
Approach: They propose a dual-encoding Transformer architecture that uses a structural encoder and a semantic encoder to seek for semantically relevant nodes.
Outcome: The proposed architecture achieves superior performance compared to state-of-the-art attention-based methods on complex relational graphs like KGs and citation networks.
On Sample Based Explanation Methods for NLP: Faithfulness, Efficiency and Semantic Evaluation (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for explaining "black-box" models such as Influence Functions are becoming more popular.
Approach: They propose a semantic-based evaluation metric that can better align with humans’ judgment of explanations than the widely adopted diagnostic or re-training measures.
Outcome: The proposed method can better align with humans’ judgment of explanations than diagnostic or re-training measures.
WARM: A Weakly (+Semi) Supervised Math Word Problem Solver (2022.coling-1)

Copied to clipboard

Challenge: Existing approaches to solving math word problems require full supervision in the form of intermediate equations.
Approach: They propose a weakly supervised model that requires only the final answer as supervision to solve math word problems.
Outcome: The proposed model achieves accuracy gains of 4.5% and 32% over current weakly-supervised methods on standard Math23K and AllArith datasets.
Detecting Cybercrimes in Accordance with Pakistani Law: Dataset and Evaluation Using PLMs (2024.lrec-main)

Copied to clipboard

Challenge: Roman Urdu is a widely used language in Pakistan but lacks sufficient resources and tools for text-based cybercrime detection.
Approach: They propose to use a benchmark dataset for text-based cybercrime detection in Roman Urdu to improve the performance of pre-trained language models.
Outcome: The proposed model achieves the highest performance on all metrics.
Named Entity Recognition via Noise Aware Training Mechanism with Data Filter (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for named entity recognition (NER) do not distinguish noisy from hard samples.
Approach: They propose a noise-aware-with-filter method to help model identify noisy samples . they propose 'incomplete trust' loss function which boosts L CRF with a robust term .
Outcome: The proposed method outperforms the existing methods on six real-world Chinese and English NER datasets.
Combinatory Grammar Tells Underlying Relevance among Entities (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches focus on dependencies among words while paying limited attention to other types of syntactic structure.
Approach: They propose an alternative approach that takes advantage of combinatory categorial grammar to detect the relation between entities.
Outcome: The proposed model performs state-of-the-art on two widely used English benchmark datasets.
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation (2025.findings-naacl)

Copied to clipboard

Challenge: Text-to-image generation models exhibit a strong bias toward English-speaking cultures, ignoring or misrepresenting the unique characteristics of other language groups, countries, and nationalities.
Approach: They propose a RusCode benchmark to evaluate the quality of text-to-image generation containing elements of the Russian cultural code.
Outcome: The proposed model is based on 1250 text prompts in Russian and their translations into English.
Attention Understands Semantic Relations (2022.lrec-1)

Copied to clipboard

Challenge: Present-day monopoly of foundation language models in most tasks forces researchers and practitioners to rely on popular large models without genuinely understanding the models' behaviour.
Approach: They propose a probing pipeline to study the representedness of semantic relations in transformer language models and propose 'attention mechanisms' that focus on syntactic relational information and semantic one.
Outcome: The proposed pipeline shows that attention scores are expressive as output activations on this task, despite their lesser ability to represent surface cues.
UniMath: A Foundational and Multimodal Mathematical Reasoner (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for interpreting and processing diverse mathematical modalities are limited . existing systems are limited in interpreting complex mathematical tasks and implementing them in a multimodal manner.
Approach: They propose a multimodal mathematical reasoning system that utilizes a fine-tuned T5 model augmented with a variational autoencoder (VAE)-based image tokenizer.
Outcome: The proposed model achieves state-of-the-art performance on SVAMP, GeoQA, and TableMWP datasets and is generalized on two additional datasets.
Contrastive Learning for Task-Independent SpeechLLM-Pretraining (2025.findings-acl)

Copied to clipboard

Challenge: Large language models excel in speech processing tasks but their reliance on written text limits their application in real-world scenarios.
Approach: They propose a task-independent speech pretraining stage and task-specific fine-tuning stage to adapt LLMs to speech processing tasks.
Outcome: The proposed model outperforms models specialized on speech translation and question answering while being trained on 10% of the task-specific data.
A Two-Agent Game for Zero-shot Relation Triplet Extraction (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for relation triplet extraction rely on labeled data and are limited in their applicability.
Approach: They propose a two-agent game approach to deliberate and debate unseen relations by two agents, a generator and an extractor.
Outcome: The proposed method outperforms baseline methods by 6%-16% in F1 scores.
LLM4RE: A Data-centric Feasibility Study for Relation Extraction (2025.coling-main)

Copied to clipboard

Challenge: Relation Extraction (RE) is a critical step in information extraction due to its wide-scale applicability for downstream applications such as Knowledge Base creation and Question Answering (QA).
Approach: They propose to conduct the first feasibility analysis to explore the viability of Large Language Models for RE by investigating their robustness to various RE scenarios stemming from data-specific characteristics.
Outcome: The proposed models are robust to various RE scenarios stemming from data-specific characteristics, but their performance is not yet fully understood.
Comparing Prompt-Based and Standard Fine-Tuning for Urdu Text Classification (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in natural language processing have demonstrated the efficacy of pre-trained language models for various downstream tasks.
Approach: They compare prompt-based fine-tuning with standard fine-uning for text classification in Urdu and Roman Urdu languages.
Outcome: The proposed approach improves up to 13% in accuracy in low-resource languages with limited labeled examples over standard fine-tuning approaches.
Copyright Violations and Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: a recent study examines the extent to which language models can memorize training data . a fair use exemption to copyright laws allows for limited use of copyrighted material .
Approach: They examine the extent to which language models can redistribute copyrighted text . they use a range of popular books and coding problems to study copyright violations .
Outcome: This study examines the extent to which language models can redistribute copyrighted text . it shows that language models may memorize entire chunks of training data .
Low-Rank Softmax Can Have Unargmaxable Classes in Theory but Rarely in Practice (2022.acl-long)

Copied to clipboard

Challenge: Probabilistic multiclass classifiers with large number of output classes are commonplace in natural language processing.
Approach: They propose to use argmax to predict words from a large vocabulary in NLP models . they find that 13 out of 150 models do indeed have such unargmaxable tokens .
Outcome: The proposed algorithms detect unargmaxable tokens in large language models and translation models.
A Statutory Article Retrieval Dataset in French (2022.acl-long)

Copied to clipboard

Challenge: Statutory article retrieval is the task of automatically retrieving law articles relevant to a legal question.
Approach: They propose to use a Belgian Statutory Article Retrieval Dataset to test various retrieval approaches including lexical and dense architectures to achieve a 74.8% R@100.
Outcome: The proposed dataset outperforms existing systems in both zero-shot and supervised setups.
Argumentation Mining on Essays at Multi Scales (2020.coling-main)

Copied to clipboard

Challenge: Argumentation mining on essays is a new task in natural language processing.
Approach: They propose a multi-scale argumentation mining model which aims to identify the types and locations of argumentation components from essay text.
Outcome: The proposed model outperforms existing models on mining all types of argumentation components on the Persuasive Essay dataset.
Modeling Conceptual Attribute Likeness and Domain Inconsistency for Metaphor Detection (2023.emnlp-main)

Copied to clipboard

Challenge: Metaphor detection aims to distinguish between metaphorical and literal expressions in text.
Approach: They propose an attribute likeness and domain inconsistency learning framework for wordpair metaphor detection based on conceptual metaphor theory . they model attribute likeity with an attribute siamese network and devise a domain contrastive learning strategy to learn semantic inconsistentness of concepts in source and target domains .
Outcome: The proposed framework outperforms existing word-pair and token-level methods on four datasets.
Probabilistic Transformer: A Probabilistic Dependency Model for Contextual Word Representation (2023.findings-acl)

Copied to clipboard

Challenge: Syntactic structures were deemed essential in natural language processing . but since the deep learning revolution, NLP has been dominated by neural models that do not consider syntactical structures in their design.
Approach: They propose a model that models latent representations of words in a sentence . they use a conditional random field to model latent and dependency arcs .
Outcome: The proposed model performs competitively to transformers on small to medium sized datasets.
EFTNAS: Searching for Efficient Language Models in First-Order Weight-Reordered Super-Networks (2024.lrec-main)

Copied to clipboard

Challenge: Depending on the size of transformer-based models, they can be restricted from deployment in resource-constrained environments.
Approach: They propose to combine neural architecture search and network pruning techniques to generate and train weight-sharing super-networks that contain efficient transformer-based models.
Outcome: The proposed model achieves high-performing, high-performance subnetworks on the general language understanding evaluation and the Stanford Question Answering Dataset.
Perspective Transition of Large Language Models for Solving Subjective Tasks (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized the field of natural language processing . performance of LLMs on subjective tasks is limited, authors say .
Approach: They propose a method that allows LLMs to select between direct, role, and third-person perspectives for best way to solve corresponding subjective problem.
Outcome: The proposed method outperforms widely used single fixed perspective based methods on 12 subjective tasks.
Emotion Analysis in NLP: Trends, Gaps and Roadmap for Future Directions (2024.lrec-main)

Copied to clipboard

Challenge: Emotion analysis (EA) is a rapidly growing field in natural language processing . there is no consensus on scope, direction, or methods for EA .
Approach: They review 154 relevant NLP papers on emotion analysis from the last decade . they ask: how are EA tasks defined in NLP? what are the most prominent emotion frameworks and which emotions are modeled?
Outcome: The authors examine 154 relevant NLP papers on emotion analysis from the last decade . they find that there is no consensus on scope, direction, or methods .
TWBias: A Benchmark for Assessing Social Bias in Traditional Chinese Large Language Models through a Taiwan Cultural Lens (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models have shown remarkable capabilities in natural language processing, but concerns about social bias amplification remain.
Approach: They propose a social bias evaluation benchmark for Traditional Chinese LLMs that integrates chat templates and diverse prompts for comprehensive bias assessment.
Outcome: The proposed model incorporates chat templates and diverse prompts for comprehensive bias assessment focusing on Taiwan's cultural context and prioritizing gender and ethnicity bias evaluation.
Annotation and Analysis of Extractive Summaries for the Kyutech Corpus (L18-1)

Copied to clipboard

Challenge: Summarization of multi-party conversation requires corpora to analyze characteristics of conversations and construct a method for summary generation.
Approach: They propose to annotate a Japanese conversation corpus for a decision-making task . they compare extractive summarization methods with the annotated extractive summary .
Outcome: The proposed corpus is the first annotated for conversation summarization tasks and freely available to anyone.
Multi-hop Graph Convolutional Network with High-order Chebyshev Approximation for Text Reasoning (2021.acl-long)

Copied to clipboard

Challenge: Existing single-hop graph reasoning in Graph convolutional networks may miss some important non-consecutive dependencies.
Approach: They propose a graph convolutional network with the high-order dynamic Chebyshev approximation which augments multi-hop graph reasoning by fusing messages aggregated from direct and long-term dependencies into one convolutionalist layer.
Outcome: The proposed model improves on four transductive and inductive NLP tasks and the ablation of the existing model.
A Topic Augmented Text Generation Model: Joint Learning of Semantics and Structural Features (D19-1)

Copied to clipboard

Challenge: Existing methods for text generation are limited in supervised setting and designed for specific applications.
Approach: They propose a text generation model that learns semantics and structural features simultaneously . their model leverages a topic-based model to enhance the recognition of text semantics .
Outcome: The proposed model outperforms state-of-the-art models in terms of text perplexity and topic coherence.
ABC: Attention with Bounded-memory Control (2022.acl-long)

Copied to clipboard

Challenge: Existing approaches to attention with bounded-memory control (ABC) have a quadratic complexity in sequence lengths, making it prohibitive for long sequences.
Approach: They propose a new abstraction that bounds memory size to improve efficiency . they propose bounded-memory control, which connects several efficient attention variants .
Outcome: The proposed approach outperforms existing approaches on language modeling, machine translation, and masked language model finetuning.
One Unified Model for Diverse Tasks: Emotion Cause Analysis via Self-Promote Cognitive Structure Modeling (2025.naacl-long)

Copied to clipboard

Challenge: Existing models for emotion cause analysis overlook common ground rooted in cognitive emotion theories, in particular, the cognitive structure of emotions.
Approach: They propose a unified model capable of tackling diverse emotion cause analysis tasks . they propose 'self-promote mechanism' that constructs the emotion cognitive structure through LLM .
Outcome: The proposed model outperforms existing models and baselines on multiple emotion cause analysis tasks.
DB-LLM: Accurate Dual-Binarization for Efficient LLMs (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for ultra-low bit quantization cause severe accuracy drops . a novel Dual-Binarization method is proposed for efficient Large Language Models .
Approach: They propose a Dual-Binarization method that takes 2-bit-width and binarization into account . they propose DB-LLM, which uses a 2-bit binarized weighted model to represent weights efficiently .
Outcome: The proposed method surpasses the current State-of-the-Art in ultra-low bit quantization and achieves 20% reduction in computational consumption compared to the SOTA method under the same bit-width.
Enhance Robustness of Language Models against Variation Attack through Graph Integration (2024.lrec-main)

Copied to clipboard

Challenge: Pre-trained language models (PLMs) are used in many NLP applications but their vulnerability to adversarial attacks can lead to false or misleading information being distributed.
Approach: They propose a method to incorporate a Chinese character variation graph into pre-trained language models to increase their robustness against character variation attacks in Chinese content.
Outcome: The proposed method outperforms existing language models in combating adversarial attacks in Chinese content.
Text Style Transferring via Adversarial Masking and Styled Filling (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models for text style transfer suffer from two challenges: the word masking procedure may mistakenly remove unexpected words and the selected words in the word filling procedure lack diversity and semantic consistency.
Approach: They propose a style transfer model with adversarial masking and styled filling techniques to solve these challenges.
Outcome: The proposed model performs well on two benchmark text style transfer data sets.
When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have achieved impressive performance across NLP tasks.
Approach: They propose to use long-context SFT to improve short-contemporary performance . they also decouple and analyze two key components, Multi-Head Attention and Feed-Forward Network .
Outcome: The proposed model improves short-context performance, contrary to pretraining.
GanLM: Encoder-Decoder Pre-training with an Auxiliary Discriminator (2023.acl-long)

Copied to clipboard

Challenge: Existing pre-training methods underutilize the benefits of language understanding for generation.
Approach: They propose a GAN-style model for encoder-decoder pre-training with an auxiliary discriminator.
Outcome: The proposed model outperforms existing pre-trained models and achieves state-of-the-art performance.
On Efficient Retrieval of Top Similarity Vectors (D19-1)

Copied to clipboard

Challenge: Existing representation learning methods such as Word2vec represent word embeddings in the semantic space.
Approach: They propose an efficient method for searching vectors via a non-metric matching function: inner product.
Outcome: Experiments on data representations learned for different machine learning tasks show the proposed method outperforms existing methods.
Visual Grounding Annotation of Recipe Flow Graph (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies have ground visual observations with procedural texts with graphs to understand which objects are aligned with textual descriptions.
Approach: They propose to provide visual grounding annotations to recipe flow graphs by adding bounding boxes to image sequences of recipes and annotating two types of event attributes with each bounding box.
Outcome: The proposed dataset gives visual grounding with workflow’s contextual information between procedural text and visual observation in an indirect manner.
Aligning Images and Text with Semantic Role Labels for Fine-Grained Cross-Modal Understanding (2022.lrec-1)

Copied to clipboard

Challenge: Currently, image retrieval systems can retrieve relevant results for diverse inputs, but they do not provide a way to intentionally inject variety into the search results.
Approach: They propose a multimodal dataset that combines semantic annotations with image bounding boxes.
Outcome: The proposed system improves image retrieval performance and flexibility.
Annotating Event Appearance for Japanese Chess Commentary Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Recent studies show that text and non-text data are not always a “true” pair.
Approach: They propose "Event Appearance" labels that show the relationship between events mentioned in texts and those happening in the real world.
Outcome: The proposed labels show the relationship between events mentioned in texts and those happening in the real world.
“That Is a Suspicious Reaction!”: Interpreting Logits Variation to Detect NLP Adversarial Attacks (2022.acl-long)

Copied to clipboard

Challenge: Existing methods to detect adversarial text inputs are limited in performance and are not detectable via spell checkers.
Approach: They propose a model-agnostic detector of adversarial text examples that detects patterns in the logits of the target classifier when perturbing the input text.
Outcome: The proposed detector improves the state-of-the-art performance in recognizing adversarial inputs and exhibits strong generalization capabilities across different NLP models, datasets, and word-level attacks.
A Measure-Theoretic Characterization of Tight Language Models (2023.acl-long)

Copied to clipboard

Challenge: Language modeling is a core task in natural language processing.
Approach: They propose to characterize leakage onto the set of infinite sequences by a measure-theoretic approach.
Outcome: The proposed language model families are tight, meaning they will not leak . the proposed language models are based on the 'sequence leakage' hypothesis .
Can Pre-trained Language Models Interpret Similes as Smart as Human? (2022.acl-long)

Copied to clipboard

Challenge: Simile interpretation is a crucial task in natural language processing.
Approach: They propose a task to let PLMs infer the shared properties of similes by probing textual corpora and human-designed questions.
Outcome: The proposed task outperforms pre-trained language models on simile interpretation tasks while still underperforming humans.
Unlocking Memorization in Large Language Models with Dynamic Soft Prompting (2024.emnlp-main)

Copied to clipboard

Challenge: Pretrained large language models excel in a variety of natural language processing tasks . however, they pose significant security risks due to their tendency to memorize training data .
Approach: They propose a method to estimate LLM memorization using dynamic, prefix-dependent soft prompts.
Outcome: The proposed method can achieve maximum relative improvement of 135.3% and 39.8% over baseline compared to state-of-the-art methods.
Generating Fluent Adversarial Examples for Natural Languages (P19-1)

Copied to clipboard

Challenge: Current methods for building adversarial attackers for NLP are inefficient as the gradient is discarded.
Approach: They propose an adversarial attacker which performs Metropolis-Hastings sampling with the guidance of gradients to solve these problems.
Outcome: The proposed algorithm outperforms the baseline model on attacking capability on IMDB and SNLI.
Sequential Learning of Convolutional Features for Effective Text Classification (D19-1)

Copied to clipboard

Challenge: Existing models for text classification have largely ignored convolution filters and max pooling . text classification is one of the major applications of natural language processing .
Approach: They propose a convolutional attentive recurrent network model which uses convolution filters and max pooling to improve text classification.
Outcome: The proposed model outperforms existing convolutional models on text classification tasks.
Sarcasm-R1: Enhancing Sarcasm Detection through Focused Reasoning (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for sarcasm detection are limited by supervised learning or prompt engineering . a new approach decomposes sarcasm detection into three dimensions: language, context, and emotion .
Approach: They propose a method that decomposes sarcasm detection into three dimensions: language, context, and emotion.
Outcome: The proposed method outperforms state-of-the-art methods in most cases.
COUNT: COntrastive UNlikelihood Text Style Transfer for Text Detoxification (2023.findings-emnlp)

Copied to clipboard

Challenge: Text detoxification is a task to ensure the generation of non-toxic and safe text.
Approach: They propose a novel contrastive unlikelihood objective that combines rephrasing and identity mapping to effectively isolate and focus learning on non-toxic style transfer.
Outcome: The proposed method achieves significant improvements in fluency, content preservation, and detoxification on two parallel datasets.
SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities (2022.acl-long)

Copied to clipboard

Challenge: Existing evaluation methods for transfer learning are limited in speech research . authors show that pre-trained models transfer well across multiple tasks .
Approach: They propose a benchmark to evaluate pre-trained models by increasing task diversity and difficulty over SUPERB.
Outcome: The proposed benchmark increases task diversity and difficulty over SUPERB-SG.
French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English (2022.acl-long)

Copied to clipboard

Challenge: We introduce 1,679 sentence pairs in French that cover stereotypes in ten types of bias like gender and age.
Approach: They build on the US-centered CrowS-pairs dataset to create a multilingual stereotypes dataset that allows for comparability across languages and cultures.
Outcome: The proposed dataset allows for comparability across languages while characterizing biases that are specific to each country and language.
Merge and Label: A Novel Neural Network Architecture for Nested NER (P19-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is one of the best studied tasks in natural language processing.
Approach: They propose a neural network architecture that merges tokens and/or entities into nested entities and labels them independently.
Outcome: The proposed approach achieves state-of-the-art F1 of 74.6 and improves with contextual embeddings to 82.4.
Annotated Corpus of Scientific Conference’s Homepages for Information Extraction (L18-1)

Copied to clipboard

Challenge: a corpus of scientific conferences contains homepages with annotations of important information . name of conference, abbreviation, place, submission, notification, camera ready dates are included .
Approach: They propose a corpus that contains 943 homepages of scientific conferences with annotations of interesting information.
Outcome: The proposed corpus contains 943 homepages of scientific conferences, 14794 including subpages . the results show that it can be used as a reference data set for this type of task.
Analyzing and Improving Coherence of Large Language Models in Question Answering (2025.naacl-long)

Copied to clipboard

Challenge: Large language models suffer from instability or lack of coherence when receiving diverse input variations.
Approach: They analyze the behavior of large language models when dealing with multiple lexical variations of the same info-seeking questions.
Outcome: The proposed model generates equivalent outputs when receiving diverse input variations.
Large Language Models are biased to overestimate profoundness (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in natural language processing have been suggested to approach AI . however, it is still unclear whether LLMs possess similar reasoning abilities to humans .
Approach: They evaluate GPT-4 and other LLMs in judging the profoundness of mundane statements . they find a significant correlation between the LLM and humans .
Outcome: The proposed model overestimates the profoundness of nonsensical statements . the model overstates the profound nature of non-senior statements, the study finds .
TURNA: A Turkish Encoder-Decoder Language Model for Enhanced Understanding and Generation (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in natural language processing have favored well-resourced English-centric models, resulting in a significant gap with low-resource languages.
Approach: They propose a language model for the low-resource language Turkish that is capable of both natural language understanding and generation tasks.
Outcome: The proposed model outperforms multilingual models in understanding and generation tasks and competes with monolingual models for understanding tasks.
DeTiME: Diffusion-Enhanced Topic Modeling using Encoder-decoder based LLM (2023.findings-emnlp)

Copied to clipboard

Challenge: Neural Topic Models and Large Language Models (LLMs) primarily use contextual embeddings from LLMs, which are not optimal for clustering or topic generation.
Approach: They propose a framework that leverages Encoder-Decoders to generate highly clusterable embeddings that could generate topics that exhibit enhanced clusterability and enhanced semantic coherence compared to existing methods.
Outcome: The proposed framework is efficient to train and exhibits high adaptability, demonstrating its potential for a wide array of applications.
Counterfactuals of Counterfactuals: a back-translation-inspired approach to analyse counterfactual editors (2023.findings-acl)

Copied to clipboard

Challenge: Existing explanations for classifiers are counterfactual or contrastive . lack of universal ground truth for counterf actual edits hinders their evaluation .
Approach: They propose a back translation-inspired evaluation methodology that utilises earlier outputs of the explainer as ground truth proxies to investigate the consistency of explainers.
Outcome: The proposed method can provide valuable insights into the behaviour of predictor and explainer models and infer patterns that would otherwise be obscured.
Part-of-Speech Tagging for Arabic Gulf Dialect Using Bi-LSTM (L18-1)

Copied to clipboard

Challenge: Part-of-speech (POS) tagging is one of the most important building blocks in many natural language processing (NLP) applications.
Approach: They propose to use a POS tagger for Arabic Gulf dialect to improve POS tagging accuracy.
Outcome: The proposed POS tagger improves POS tagging accuracy for the Arabic Gulf dialect from 75% accuracy to 91% accuracy using a bi-LSTM labeler.
PatchBERT: Just-in-Time, Out-of-Vocabulary Patching (2020.emnlp-main)

Copied to clipboard

Challenge: a pre-trained language model with low OOV can improve performance for transfer learning . a vocabulary surrogate can provide performance boosts with no additional computation cost .
Approach: They propose multiple methods to mitigate OOV during downstream task fine-tuning . they demonstrate that vocabulary surrogates can provide performance boosts with no additional computation cost .
Outcome: The proposed methods improve performance with the same parameter count when combined with fine-tuning.
Decomposition for Enhancing Attention: Improving LLM-based Text-to-SQL through Workflow Paradigm (2024.findings-acl)

Copied to clipboard

Challenge: In-context learning of large-language models has achieved remarkable success in the field of natural language processing . however, the single-step chain-of-thought prompting approach faces challenges such as attention diffusion and inadequate performance in complex tasks like text-to-SQL.
Approach: They propose a workflow paradigm method to enhance the attention and problem-solving scope of large-language models through decomposition.
Outcome: The proposed method outperforms existing methods on three datasets and improves the upper limit of LLM-based approaches.
Deep Attention Diffusion Graph Neural Networks for Text Classification (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for text classification based on graph neural networks (GNNs) consider only one-hop neighborhoods and low-frequency information within texts, which suffer from over-smoothing issues if many graph layers are stacked.
Approach: They propose a deep attention diffusion Graph Neural Network model to learn text representations by bridging the chasm of interaction difficulties between a word and its distant neighbors.
Outcome: The proposed model outperforms existing methods on standard benchmark datasets on a set of textual features.
Balancing Methods for Multi-label Text Classification with Long-Tailed Class Distribution (2021.emnlp-main)

Copied to clipboard

Challenge: Multi-label text classification is a challenging task because it requires capturing label dependencies.
Approach: They propose to use distribution-balanced loss functions to solve label dependency problems in multi-label text classification by capturing label dependencies from a fixed-set of labels.
Outcome: The proposed loss function addresses both the class imbalance and label linkage problems and outperforms other loss functions.
GRAF: Graph Retrieval Augmented by Facts for Romanian Legal Multi-Choice Question Answering (2025.findings-acl)

Copied to clipboard

Challenge: Question answering systems have been used for various domains and languages.
Approach: They propose a novel approach for question answering (QA) that combines a dataset of Romanian legal questions with a CROL corpus of laws.
Outcome: The proposed approach achieves competitive results with generally accepted state-of-the-art methods and even exceeds them in most settings.
SenSALDO: Creating a Sentiment Lexicon for Swedish (L18-1)

Copied to clipboard

Challenge: sentiment analysis has seen an explosive expansion over the last decade or so . many theoretical and methodological questions remain unanswered and resource gaps unfilled .
Approach: They develop a sentiment lexicon for written (standard) Swedish using an existing dataset . they assign a real value sentiment score in the range [-1,1] and produce a label for it .
Outcome: The proposed sentiment lexicon is an open source resource from the Swedish Language Bank . it is based on an existing gold standard dataset and is available from Sprkbanken .
Improving Image Captioning with Better Use of Caption (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to image captioning focus on visual attention, but many do not.
Approach: They propose a framework that explores semantics available in captions and leverages that to enhance both image representation and caption generation.
Outcome: The proposed framework outperforms baselines on the MSCOCO dataset and is state-of-the-art under a wide range of evaluation metrics.
Schema Generation for Large Knowledge Graphs Using Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Schemas are a vital part of ontology engineering and require substantial knowledge engineers and domain experts to create them.
Approach: They propose to use large language models to generate schemas in Shape Expressions (ShEx) to bridge the resource gap between knowledge engineers and domain experts.
Outcome: The proposed pipelines use local and global information from knowledge graphs (KGs) to generate high-quality schemas in Shape Expressions (ShEx).
Linghub2: Language Resource Discovery Tool for Language Technologies (2022.lrec-1)

Copied to clipboard

Challenge: Linghub is a platform for language resources that can be used to find and retrieve data . the platform is based on a popular open source data management system, DSpace .
Approach: This work describes a rejuvenation and modernisation of the 2015 platform into using a popular open source data management system, DSpace, as foundation.
Outcome: Linghub2 1 aims to help language resources and technology users find and retrieve relevant data . the new platform, Ling hub2, contains updated and extended resources and more languages offered .
Reproducing Neural Ensemble Classifier for Semantic Relation Extraction inScientific Papers (2020.lrec-1)

Copied to clipboard

Challenge: Replicability and reproducibility are core ideas of modern scientific methods.
Approach: They describe challenges encountered in reproducing the results of a top performing system in computational linguistics.
Outcome: The proposed system was able to reproduce the results of a task 7 in the domain of natural language processing and computational linguistics.
CReTIHC: Designing Causal Reasoning Tasks about Temporal Interventions and Hallucinated Confoundings (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated impressive capabilities in natural language processing, but their ability to establish causal relationships remains challenging.
Approach: They propose a novel dataset to test and enhance the causal reasoning abilities of large language models (LLMs) by integrating elements of verbal hallucinations and temporal interventions into existing causal inference datasets.
Outcome: The proposed dataset is designed to test and enhance the causal reasoning abilities of large language models.
Resilience of Large Language Models for Noisy Instructions (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are powerful tools for interpreting human commands and generating text.
Approach: They examine the resilience of large language models against five common types of disruptions including ASR, OCR, grammatical errors, typographical errors and distractive content.
Outcome: The models show resistance to noise, but their performance suffers . authors evaluated the models against five common types of disruptions based on their results .
GeezSwitch: Language Identification in Typologically Related Low-resourced East African Languages (2022.lrec-1)

Copied to clipboard

Challenge: Low-resourced languages with similar typologies are often confused with each other in real-world applications such as machine translation, affecting the user’s experience.
Approach: They propose to build a dataset for five typologically and phylogenetically related low-resourced East African languages using the Ge’ez script as a writing system.
Outcome: The proposed dataset is built automatically from selected data sources, but also performed a manual evaluation to assess its quality.
Lost in Translation: Benchmarking Commercial Machine Translation Models for Dyslexic-Style Text (2025.findings-acl)

Copied to clipboard

Challenge: Dyslexia affects writing, leading to unique patterns such as letter and homophone swapping.
Approach: They examine the fairness of four commercial machine translation systems towards dyslexic text through a systematic audit using both synthetically generated and real writing from individuals with dyslexia.
Outcome: The proposed system audits show that it is fair to use synthetic and synthetic dyslexic text and real writing from people with dyslexia.
Attention-Enhancing Backdoor Attacks Against BERT-based Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing textual backdoor attacks focus on generating stealthy triggers or modifying model weights.
Approach: They propose a Trojan Attention Loss (TAL) which enhances the Trojan behavior by directly manipulating attention patterns.
Outcome: The proposed method improves the effectiveness of the backdoor attacks on different backbone models and tasks.
How Speculative Can Speculative Decoding Be? (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) have a largely increased latency due to their ability to autoregressively model . speculative decoding is a technique that trades generation quality for speed .
Approach: They propose to use a draft model to draft tokens autoregressively and then verify them in parallel.
Outcome: The proposed model could draft tokens autoregressively and then verify them in parallel . the proposed model trades quality for speed and could fail in verification stage .
On the Sentence Embeddings from Pre-trained Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Pre-trained contextual representations like BERT have been widely used for NLP tasks.
Approach: They propose to transform anisotropic sentence embedding distribution to smooth and isotropic Gaussian distribution by normalizing flows that are learned with an unsupervised objective.
Outcome: The proposed method achieves significant performance gains over state-of-the-art embeddings on a variety of semantic textual similarity tasks.
Ontology-Guided Reverse Thinking Makes Large Language Models Stronger on Knowledge Graph Question Answering (2025.acl-long)

Copied to clipboard

Challenge: Existing methods rely on entity vector matching, but the purpose of the question is abstract and difficult to match with specific entities. Existing approaches rely only on entity-vector matching, and there is a problem with multi-hop reasoning.
Approach: They propose a framework that constructs reasoning paths from purposes back to conditions using the KG ontology.
Outcome: Experiments on the WebQSP and CWQ datasets show that ORT significantly improves the capability of large language models in knowledge graph question answering tasks (KGQA).
LDIR: Low-Dimensional Dense and Interpretable Text Embeddings with Relative Representations (2025.findings-acl)

Copied to clipboard

Challenge: Existing text embeddings with high dimensions are difficult to trace and interpret.
Approach: They propose low-dimensional and interpretable text embeddings with relative representations that encode semantic meanings in a vector space where similar texts are close together in the representation space.
Outcome: The proposed embeddings outperform existing models on multiple tasks with fewer dimensions and are lowdimensional and dense while maintaining interpretability.
Predicting Performance for Natural Language Processing Tasks (2020.acl-main)

Copied to clipboard

Challenge: Natural language processing (NLP) is a vast field, with a wide variety of tasks, languages, and domains.
Approach: They build regression models to predict evaluation score of an NLP experiment . they find that their models can produce meaningful predictions over unseen languages .
Outcome: The proposed model outperforms baseline models and human experts on 9 different tasks.
Eyes Show the Way: Modelling Gaze Behaviour for Hallucination Detection (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for hallucination detection depend on knowledge sources that are explicit such as Wikipedia or knowledge graphs.
Approach: They propose a cognitive approach that leverages gaze signals from humans to detect hallucinations in natural language processing (NLP) they collect and introduce an eye tracking corpus consisting of 500 instances, annotated by five annotators for hallucinism detection.
Outcome: The proposed approach achieves a balanced accuracy of 87.1% on a FactCC dataset.
Can LLMs Understand the Implication of Emphasized Sentences in Dialogue? (2024.findings-emnlp)

Copied to clipboard

Challenge: Emphasis is a crucial component in human communication, which indicates speaker’s intention and implication beyond pure text in dialogue.
Approach: They propose a benchmark dataset with annotated dialogue samples capturing the implications of emphasis.
Outcome: The proposed evaluation pipeline achieves high correlation with human scoring and commercial LLMs perform better than open-source LLM.
SC-CoMIcs: A Superconductivity Corpus for Materials Informatics (2020.lrec-1)

Copied to clipboard

Challenge: Existing corpus of superconducting materials in Materials Informatics (MI) is limited.
Approach: They propose to create a corpus tailored for the text mining of superconducting materials in Materials Informatics.
Outcome: The proposed corpus can find terms relevant to a query term within a specified Named Entity category.
Conversation Chronicles: Towards Diverse Temporal and Relational Dynamics in Multi-Session Conversations (2023.emnlp-main)

Copied to clipboard

Challenge: open-domain chatbots focus on short single-session dialogue, neglecting the potential need for understanding contextual information in multiple consecutive sessions.
Approach: They propose a 1M multi-session dialogue dataset for integrating time intervals and speaker relationships into a long-term conversation setup.
Outcome: The proposed model can generate coherent responses according to time intervals and speaker relationships with high user engagement without contradiction in a long-term conversation setup.
Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have advanced natural language processing, but their effectiveness is often hampered by parameter mis-filling during tool calling.
Approach: They propose a hierarchical tool error checklist framework to diagnose and mitigate tool-calling errors without relying on extensive real-world interactions.
Outcome: The proposed framework improves parameter-filling accuracy and tool-calling success rates compared to baseline methods.
Learning Event-aware Measures for Event Coreference Resolution (2023.findings-acl)

Copied to clipboard

Challenge: Existing models for event coreference resolution are based on entity-level tasks, but event coreferent resolution is a challenge.
Approach: They propose a model that learns and integrates multiple representations from event alone and event pair on the basis of event but not entity as before.
Outcome: The proposed model achieves new state-of-the-art on the ACE 2005 benchmark, demonstrating the effectiveness of the proposed framework.
End-to-end Aspect-based Sentiment Analysis with Combinatory Categorial Grammar (2023.findings-acl)

Copied to clipboard

Challenge: End-to-end aspect-based sentiment analysis (EASA) is a natural language processing task that requires a deep understanding of the running text.
Approach: They propose a method to improve EASA with CCG supertags that carry syntactic and semantic information of the associated words.
Outcome: The proposed approach outperforms baselines and achieves state-of-the-art results on all datasets.
LLMaAA: Making Large Language Models as Active Annotators (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing supervised learning methods in natural language processing require large amounts of data.
Approach: They propose an active learning loop that takes LLMs as annotators and puts them into an active loop to determine what to annotate efficiently.
Outcome: The proposed model outperforms existing models with few-shot performance in two NLP tasks.
Sparsity-Accelerated Training for Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated proficiency across various NLP tasks but often require additional training, such as continual pre-training and supervised fine-tuning.
Approach: They propose to leverage sparsity in pre-trained LLMs to accelerate training by disregarding computations for unimportant neurons.
Outcome: The proposed framework achieves comparable or superior performance to standard training while significantly accelerating the process.
Text Fluoroscopy: Detecting LLM-Generated Text through Intrinsic Features (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized the field of natural language processing because of their excellent performance on various tasks.
Approach: They propose a black-box method with better generalizability for detecting LLM-generated text by mining the intrinsic features of the text to be detected.
Outcome: The proposed method achieves 7.36% and 2.84% improvement in detection performance compared to baselines in detecting texts from different domains generated by GPT-4 and Claude3 respectively.
A Tree Extension for CoNLL-RDF (2020.lrec-1)

Copied to clipboard

Challenge: CoNLL-RDF provides a bridge for popular oneword-per-line formats . main reasons for their popularity are the simplicity of tables and tab-separated values .
Approach: They propose a technology that provides a bridge between knowledge graphs and natural language processing.
Outcome: The proposed technology provides a bridge for popular one-word-per-line formats . it provides native support for word-level annotations, but not phrase structures or text structure .
Large Language Models for Generative Recommendation: A Survey and Visionary Discussions (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized the field of natural language processing but are not fully able to leverage the generative power of LLM.
Approach: They examine the progress, methods, and future directions of large language models . they examine what generative recommendation is, why RS should advance to generative recommendations .
Outcome: The proposed approach can be simplified to generate recommendations from the entire pool of items.
Interchange Formats for Visualization: LIF and MMIF (2020.lrec-1)

Copied to clipboard

Challenge: In this paper, we discuss the enhanced data visualization capabilities enabled by interoperating computational linguistics and natural language processing (NLP) applications.
Approach: They propose to use interchange formats to enable enhanced data visualization . they propose to combine CL tools with openly available visualization tools .
Outcome: The proposed formats can be used to create visualizations and manipulate annotations in multiple ways.
Here’s a Free Lunch: Sanitizing Backdoored Models with Model Merge (2024.findings-acl)

Copied to clipboard

Challenge: democratization of pre-trained language models brings significant security risks, including backdoor attacks.
Approach: They propose to merge a backdoored model with other homogeneous models to remediate backdoor vulnerabilities.
Outcome: The proposed model merging approach outperforms other models on classification tasks without additional resources or specific knowledge.
Defense Against Prompt Injection Attack by Leveraging Attack Techniques (2025.acl-long)

Copied to clipboard

Challenge: Recent attacks leverage LLMs’ instruction-following abilities and their inabilities to distinguish instructions injected in the data content.
Approach: They invert the intention of prompt injection methods to develop novel defense methods based on previous training-free attack methods by repeating the attack process with the original input instruction rather than the injected instruction.
Outcome: The proposed methods outperform existing defense approaches, achieving state-of-the-art results.
Unveiling the Essence of Poetry: Introducing a Comprehensive Dataset and Benchmark for Poem Summarization (2023.emnlp-main)

Copied to clipboard

Challenge: Summarization of poetry is a challenging task as it can be easily lost if only the literal meaning is considered.
Approach: They propose to use poetry as a model to summarize poetry and provide a dataset to evaluate their creative language interpretation capacity.
Outcome: The proposed dataset consisting of 3011 samples and its corresponding summarized interpretation in the English language provides an opportunity to evaluate the creative language interpretation capacity of the proposed models.
Explicit Memory Learning with Expectation Maximization (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models lack reliable learning mechanisms for updating information across interactions.
Approach: They propose a framework that enhances explicit memory updates via the Expectation-Maximization algorithm.
Outcome: The proposed framework outperforms existing methods without memory or with static external memory on streaming inference tasks.
RWKV: Reinventing RNNs for the Transformer Era (2023.findings-emnlp)

Copied to clipboard

Challenge: recurrent neural networks struggle to match the performance of Transformers due to limitations in parallelization and scalability.
Approach: They propose a model architecture that combines the efficient parallelizable training of transformers with the efficient inference of RNNs.
Outcome: The proposed model performs on par with similarly sized RNNs, suggesting future work can leverage this architecture to create more efficient models.
Improving Multi-hop Logical Reasoning in Knowledge Graphs with Context-Aware Query Representation Learning (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods rely on linear sequential operations to solve First-Order Logic queries.
Approach: They propose a model-agnostic approach that fully integrates the context of the query graph.
Outcome: The proposed method improves performance on two datasets by 19.5%.
Make Prompt-based Black-Box Tuning Colorful: Boosting Model Generalization from Three Orthogonal Perspectives (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown increasing power on NLP tasks. however, tuning these models for downstream tasks usually requires exorbitant costs.
Approach: They propose a black-box tuning technique that optimizes task-specific prompts without accessing gradients and hidden representations.
Outcome: The proposed method improves performance under few-shot learning scenarios.
Are Decoder-Only Language Models Better than Encoder-Only Language Models in Understanding Word Meaning? (2024.findings-acl)

Copied to clipboard

Challenge: Large language models are highly effective tools for solving different kinds of problems in natural language processing.
Approach: They propose to use large language models to solve a myriad of problems.
Outcome: The proposed model performs worse on word meaning comprehension than an encoder-only model with vastly fewer parameters.
Medical Entity Disambiguation with Medical Mention Relation and Fine-grained Entity Knowledge (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for medical entity disambiguation (MED) fail to fully utilize the knowledge within medical knowledge bases (KBs) Existing models overlook essential interactions between medical mentions and candidate entities, resulting in knowledge- and interaction-inefficient modeling and suboptimal disambiguations performance.
Approach: They propose to combine a mention relation fusion module and an entity knowledge fusion modules to map medical mentions to corresponding entities in a knowledge base (KB)
Outcome: The proposed method outperforms state-of-the-art MED models on two publicly available real-world datasets.
PepRec: Progressive Enhancement of Prompting for Recommendation (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have been gaining in-depth performance in natural language processing domains.
Approach: They propose a training-free prompting framework that captures knowledge from content-based filtering and collaborative filtering to boost recommendation performance with LLMs.
Outcome: The proposed framework outperforms traditional deep learning recommendation models and prompt-based recommendation systems on two real-world datasets.
ParroT: Translating during Chat using Large Language Models tuned with Human Translation and Feedback (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) like ChatGPT are only accessible through restricted APIs, which creates barriers to new research and advancements in the field.
Approach: They propose a framework to enhance and regulate the translation abilities during chat . they reformulate translation data into the instruction-following style and introduce a "Hint" field .
Outcome: The proposed framework enhances and regulates the translation abilities during chat . it reformulates translation data into the instruction-following style and introduces a "Hint" field .
NER-guided Comprehensive Hierarchy-aware Prompt Tuning for Hierarchical Text Classification (2024.lrec-main)

Copied to clipboard

Challenge: Hierarchical text classification (HTC) is a challenging task in natural language processing due to its complex taxonomic label hierarchy.
Approach: They propose to use prompts to model hierarchical text classification (HTC) they propose to introduce conditional random fields and Global Pointer to establish hierarchic dependencies .
Outcome: The proposed approach achieves state-of-the-art (SoTA) performance on three public datasets.
NSina: A News Corpus for Sinhala (2024.lrec-main)

Copied to clipboard

Challenge: introducing large language models (LLMs) has advanced natural language processing (NLP), but their effectiveness is largely dependent on pre-training resources.
Approach: They propose a large news corpus for Sinhala with a set of NLP tasks for the language . NSina is the largest news corpuse for Sinha, available up to date .
Outcome: The proposed model outperforms existing models in many benchmarks and outperformed previous models in high-resource languages.
Encode Errors: Representational Retrieval of In-Context Demonstrations for Multilingual Grammatical Error Correction (2025.findings-acl)

Copied to clipboard

Challenge: a novel method for encoding fine-grained error patterns improves performance on GEC.
Approach: They propose a method for encoding grammatical errors from LLMs' internal states using a GER method.
Outcome: The proposed method significantly boosts performance in ICL settings on multilingual GEC datasets.
PPORTAL_ner: An Annotated Corpus of Portuguese Literary Entities (2024.lrec-main)

Copied to clipboard

Challenge: Annotated corpus of 25 literary texts provides a rich set of annotations for Named Entity Recognition models.
Approach: They propose an annotation dataset that simplifies the development of Named Entity Recognition models for Portuguese literary texts.
Outcome: The proposed dataset simplifies the development of Named Entity Recognition models for Portuguese literary works.
Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarks (2025.findings-emnlp)

Copied to clipboard

Challenge: Previous work has shown degraded performance of large language models for under-represented English dialects.
Approach: They analyze the effects of typifying “standard” American English language questions as non-”standard” dialectal variants on multiple choice questions.
Outcome: The results show that typifying “standard” American English language questions as non-”standard” dialectal variants can reduce performance 20% .
Rule-Guided Extraction: A Hierarchical Rule Optimization Framework for Document-Level Event Argument Extraction (2025.findings-emnlp)

Copied to clipboard

Challenge: Document-level event argument extraction (EAE) is a critical task in natural language processing.
Approach: They propose an LLM-driven HiErarchical Rule Optimization framework that iteratively generates and selects optimal hierarchical rules.
Outcome: The proposed framework outperforms few-shot supervised methods and outperformed state-of-the-art prompting baselines.
CausalRAG: Integrating Causal Graphs into Retrieval-Augmented Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing RAG frameworks face critical limitations due to text chunking and semantic similarity.
Approach: They propose a framework that incorporates causal graphs into the retrieval process.
Outcome: The proposed framework preserves contextual continuity and improves retrieval precision, leading to more accurate and interpretable responses.
KoACD: The First Korean Adolescent Dataset for Cognitive Distortion Analysis via Role-Switching Multi-LLM Negotiation (2025.findings-emnlp)

Copied to clipboard

Challenge: Cognitive distortions are negative thinking patterns that can lead to mental health issues in adolescents.
Approach: They propose a multi-Large Language Model negotiation method to refine distortion classification . they also use cognitive clarification and cognitive balancing to improve label consistency .
Outcome: The proposed dataset contains 108,717 instances of cognitive distortions in Korean adolescents.
RevMUX: Data Multiplexing with Reversible Adapters for Efficient LLM Batch Inference (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have brought a great breakthrough to the natural language processing community, but their high throughput demands make them difficult to handle concurrent queries.
Approach: They propose a parameter-efficient data multiplexing framework that integrates a reversible design in the multiplexer and can be reused to perform reverse operations and restore individual samples for classification.
Outcome: The proposed framework improves inference efficiency while maintaining satisfactory classification performance.
CSTree-SRI: Introspection-Driven Cognitive Semantic Tree for Multi-Turn Question Answering over Extra-Long Contexts (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved remarkable success in natural language processing (NLP), particularly in single-turn question answering (QA) on short-text.
Approach: They propose a framework that captures logical correlations across chunks of ELC and maintains coherence of multi-turn Questions.
Outcome: The proposed framework is able to capture logical correlations across chunks of ELC and maintain coherence of multi-turn Questions.
RAC: Efficient LLM Factuality Correction with Retrieval Augmentation (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit impressive results across a wide range of tasks, yet they can often produce factually incorrect outputs.
Approach: They propose a low-latency post-correction method that decomposes the LLM’s output into atomic facts and applies a fine-grained verification and correction process with retrieved content to verify and correct the Llm-generated output.
Outcome: The proposed method has greatly reduced latency and token consumption up to 7x compared to previous state-of-the-art methods with similar or better performance.
Pre-trained Models Perform the Best When Token Distributions Follow Zipf’s Law (2025.emnlp-main)

Copied to clipboard

Challenge: Existing large language models typically fix a vocabulary size in advance, then use Byte Pair Encoding (BPE) to construct the tokenizer.
Approach: They propose a method for determining the vocabulary size by analyzing token frequency distributions through Zipf’s law and propose to use it to optimize model performance.
Outcome: The proposed method improves model efficiency and effectiveness across NLP, genomics, and chemistry.
SubLIME: Subset Selection via Rank Correlation Prediction for Data-Efficient LLM Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Large language models and datasets have made benchmark evaluations computationally prohibitive.
Approach: They propose a framework that reduces evaluation costs by 80% to 99% while preserving ranking fidelity.
Outcome: The proposed evaluation reduces evaluation costs by 80% to 99% while preserving ranking fidelity.
Vygotsky Distance: Measure for Benchmark Task Similarity (2024.lrec-main)

Copied to clipboard

Challenge: GLUE, SuperGLUE and RussianSuperGLUE benchmarks are arbitrary sets of tasks that are not generalized.
Approach: They propose a theoretical instrument and an algorithm to calculate similarity between benchmark tasks . they use relative performance of the "students" on a given task to determine similarity .
Outcome: The proposed model reduces the number of evaluation tasks while maintaining high validation quality.
What Is Needed for Intra-document Disambiguation of Math Identifiers? (2024.lrec-main)

Copied to clipboard

Challenge: Ambiguity in math identifiers within a document poses significant challenges to understanding formulae . ambiguity in mathematical expressions can be difficult to disambiguate, requiring intra-document disambiguation .
Approach: They propose to use position data and local formula structure to disambiguate math identifiers . they train a model that performs similarly to humans with an 85% accuracy .
Outcome: The proposed model outperforms rule-based models in natural language processing.
WikiSplit++: Easy Data Refinement for Split and Rephrase (2024.lrec-main)

Copied to clipboard

Challenge: Existing text simplification methods rely on encoder-decoder models to achieve this task.
Approach: They propose a text-to-text generation approach that applies encoder-decoder models to a large-scale dataset to improve Split and Rephrase.
Outcome: The proposed approach improves Split and Rephrase readability and performance on large datasets, but still suffers from hallucinations and under-splitting.
A Survey of Toxicity Mitigation Strategies for Multilingual Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models can reproduce and amplify toxic content, including hate speech, harassment, and bias.
Approach: They propose a comprehensive survey of the many detoxification methods tailored to multilingual LLMs.
Outcome: The proposed methods are based on data filtering, style transfer, expert-based logit steering, retrieval augmentation, and human feedback.
TinyAttack: Exploring Stylistic Vulnerabilities in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing research on robustness of large language models has focused on text-based perturbations and the use of invisible characters and homoglyphs.
Approach: They propose a framework to exploit weaknesses in Large Language Models (LLMs) by changing their stylistic structure using Unicode.
Outcome: The proposed framework exploits vulnerabilities in large language models through Unicode-based stylistic transformations without altering its semantic or syntactic structure.
Text Embedding as Treatment: A Meta Causal Approach for Robust Sentiment Classification (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for sentiment classification use binary treatment of words . Existing approaches limit generalizability to novel words and low-frequency words if there is a word in a sentence that is not treated .
Approach: They propose a meta-causal approach that uses a single training task to identify causal words for arbitrary words.
Outcome: The proposed method reduces the spurious correlation between word treatment and sentiment classification by removing words with low treatment effects from a pre-trained language model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations